How it works

Send an export. Get your first graph.

No SDK, no routing change, no security review to find out whether this works on your traffic.

You

Send past runs

Inputs, steps, tool calls, outputs, and any outcomes you track.

Us

Build the baseline

Runs clustered by intent, so comparison happens against similar work.

Us

Flag and diagnose

What stopped matching, where it began, and which lever addresses it.

You

Rule on it

Confirm, dismiss, or correct. Every verdict goes back into the graph.

Signals

Five checks. Weighted by how much graph you have.

In a dense cluster, deviation carries the call. In a sparse one, only the absolute signals count.

StructuralWorks at n=1

Did it do what it claimed?

The output says the record was updated. There is no write in the trajectory.

ContextWorks at n=1

Was the material trustworthy?

Empty retrieval that still produced a confident paragraph. An instruction evicted before the final step.

BehavioralNeeds a cluster

Did it behave like its neighbors?

A step skipped that almost every comparable run takes. A path nothing else in the cluster took.

RobustnessFlagged only

Does the answer hold?

Replay with a different seed, model, or context order. Divergence means nothing determined it.

OutcomeStrongest label

Did your systems agree?

The draft a human edited. The ticket that reopened. The intervention your team already made.

Rare doesn't mean wrong. A novel edge case is scored on what's absolute about it, and once your team says it's fine, it becomes a known-good variant.

Diagnosis

Four levers, and the evidence behind the call.

Context

Right material?

Fix: retrieval scope, injection order, or what carries between steps.

Prompt

Still doing its job?

Fix: re-engineer for the model actually serving this step, and drop dead tokens.

Model

Right model here?

Fix: upgrade steps that regressed, downgrade steps that never needed frontier.

Step design

Needs a variation?

Fix: a scoped variation for the edge case, instead of a hand-patch.

A wrong flag is a shrug. A wrong diagnosis costs an engineer a day. So every recommendation ships with the step, the context at that step, and the runs in its cluster that behaved differently.
Our bar for shipping a recommendation

Send us one week of past runs.

We'll come back with what your successful executions have in common, the few that stopped matching, and what we think changed. If we find nothing, we'll tell you that too.

Start a design partner conversation