Sample report
This is what comes back after one week of runs.
A real report, redacted and renamed. Twenty runs ranked, each with a named cause, and one worked through end to end.
Page 1 · Summary
Runs analyzed across one agent, seven days.
Intent clusters found. Nine of them sparse, scored on absolute signals only.
Runs surfaced for review. Eleven confirmed wrong by the customer's team.
Precision on this report: 11 of 20. We publish that number whether or not it flatters us.
Page 2 · The ranked list
Page 3 · One run, worked through
run 4812 — reported a write that never happened
The agent told the member their record was updated. Every check passed and the eval scored 0.91.
Why we called it context
The 1,240 comparable runs in this cluster retrieved between three and six documents at step 22. This one retrieved none and continued anyway.
Widen retrieval scope, and fail the step on empty
Two changes: broaden the query at step 22, and stop the run rather than narrating over an empty result.
Page 4 · What we couldn't tell you
- Nine sparse clusters had too few runs for behavioral comparison. We scored them on structural and context signals only.
- No downstream outcomes were supplied, so we couldn't confirm findings against retries or escalations. That's the single biggest improvement available.
- Nine of twenty flags were dismissed. Four were legitimate edge cases we now treat as known-good.
Every report ends with this page. If we can't tell you something, you should hear it from us rather than discover it.
Send us one week of past runs.
We'll come back with what your successful executions have in common, the few that stopped matching, and what we think changed. If we find nothing, we'll tell you that too.
Start a design partner conversation