- contact@insightlense.com
Design guides on agent observability and accountability, and field notes from Vsure — a specialty P&C carrier whose staff are AI agents, which we run to find out where our own products break. Every number below comes from its ledger.
How the platform is put together, and the ideas it is built on. Applicable whatever your agents actually do.
An agent that cannot be refused is not autonomous — it is unguarded. Why authority has to live in the product rather than the prompt, and what you lose when you throw refusals away.
An agent produces three different kinds of evidence, and three different people come looking for them. Most tools keep only the first.
Standard OpenTelemetry ingest, current GenAI semantic conventions and the legacy attributes people actually have in production — with no application change.
Most cost dashboards are a second source of truth that quietly disagrees with the invoice. Here is the alternative — and the metric nobody else reports.
Separate tables for human feedback and automated evals is the reason nobody ever compares them. One shape, four sources, attachable to anything.
Every trace records the prompt version that produced it, so v1 and v2 can be compared on outcomes rather than on vibes — even across two products.
An autonomy ladder is worthless if it lives in a system prompt. The mechanic that makes it real is a permission check that returns a downgrade.
Defect classes we hit in our own products, and the design decisions that came out of them.
We found the same bug five times in our own products. Every time it had passed code review, and every time the schema looked perfect.
A renewal check produced exactly the right number and then let the policy through anyway. The most dangerous control is the one that looks green.
How the pieces fit together in a real operating model, and what people outside engineering ask about agent systems.
The ten components of a property and casualty carrier, the agent workstations that sit on them, and a full worked day: a renewals alert at 08:40 followed through ops triage, a developer, an auditor and compliance — six roles, one set of records.
The measure of an observability layer is not how much it collects. It is whether these five can each be answered with one query.
Actor kind belongs in a column, not buried in a JSON payload. Otherwise the first question a regulator asks becomes a two-week project.
They are genuinely excellent at the model call. The gap opens the moment your agent is allowed to write to a system of record.
Vsure is a specialty P&C carrier whose staff are agents. These posts are what happened when it ran. Every number is drawn from its ledger.
240 insureds. $243M written. 775 claims. A live reinsurance programme. Every function run end to end by agents working through the same APIs a person uses.
One identifier, five services, fifteen steps. The whole thread of a broker submission from arrival to underwriter queue, with nothing left out.
Three catastrophes in sequence. The first ceded to treaty exactly. The third found the layer exhausted, retained $13M — and wrote down why.
Four money defects that surfaced when agents were allowed to run claims end to end — and what each one looked like in the ledger.