How it works
Send an export. Get your first graph.
No SDK, no routing change, no security review to find out whether this works on your traffic.
Send past runs
Inputs, steps, tool calls, outputs, and any outcomes you track.
Build the baseline
Runs clustered by intent, so comparison happens against similar work.
Flag and diagnose
What stopped matching, where it began, and which lever addresses it.
Rule on it
Confirm, dismiss, or correct. Every verdict goes back into the graph.
Signals
Five checks. Weighted by how much graph you have.
In a dense cluster, deviation carries the call. In a sparse one, only the absolute signals count.
Did it do what it claimed?
The output says the record was updated. There is no write in the trajectory.
Was the material trustworthy?
Empty retrieval that still produced a confident paragraph. An instruction evicted before the final step.
Did it behave like its neighbors?
A step skipped that almost every comparable run takes. A path nothing else in the cluster took.
Does the answer hold?
Replay with a different seed, model, or context order. Divergence means nothing determined it.
Did your systems agree?
The draft a human edited. The ticket that reopened. The intervention your team already made.
Rare doesn't mean wrong. A novel edge case is scored on what's absolute about it, and once your team says it's fine, it becomes a known-good variant.
Diagnosis
Four levers, and the evidence behind the call.
Right material?
Fix: retrieval scope, injection order, or what carries between steps.
Still doing its job?
Fix: re-engineer for the model actually serving this step, and drop dead tokens.
Right model here?
Fix: upgrade steps that regressed, downgrade steps that never needed frontier.
Needs a variation?
Fix: a scoped variation for the edge case, instead of a hand-patch.
A wrong flag is a shrug. A wrong diagnosis costs an engineer a day. So every recommendation ships with the step, the context at that step, and the runs in its cluster that behaved differently.
Send us one week of past runs.
We'll come back with what your successful executions have in common, the few that stopped matching, and what we think changed. If we find nothing, we'll tell you that too.
Start a design partner conversation