JournalAgentic AI

Agentic AI and the Compounding Cost of a Wrong Turn

Per-step accuracy is a seductive metric. At 95% per step, a twenty-step task succeeds only about a third of the time.

30 May 20269 min readPlanning · Reliability · Evaluation

The pitch for agentic systems is that the model stops answering and starts doing. The engineering reality is that autonomy converts a single prediction problem into a sequential one, and sequential problems multiply their errors.

The arithmetic nobody puts on the slide

steps   p=0.95    p=0.99
  5     77.4%     95.1%
 10     59.9%     90.4%
 20     35.8%     81.8%
 50      7.7%     60.5%
End-to-end success at p per-step accuracy

A 95% step is an excellent model and a poor agent. This is why demos that chain five tool calls feel magical and production workflows that chain thirty feel broken — both are behaving exactly as the arithmetic predicts.

Three levers that actually move the number

  1. Shorten the chain. Every step you can replace with a deterministic function is a step that cannot fail probabilistically. Most agent traces contain more of these than teams expect.
  2. Make steps verifiable. A step whose output can be checked cheaply — schema validation, a compile, a test run, a database constraint — converts a silent failure into a retry.
  3. Checkpoint state. If step 18 fails, resuming from 17 costs one step; restarting costs eighteen and burns the user's patience.

Autonomy is not a feature you add. It is a budget you spend, and the currency is verifiability.

What to instrument

Aggregate task success hides everything you need. Log per-step outcomes and you can see which step is eating the budget:

  • Step-level success rate, not just end-to-end completion.
  • Retry depth — a step that usually succeeds on attempt three is a step that is failing.
  • Divergence: how far the executed plan drifted from the plan proposed at step zero.
  • Human interventions per completed task. This is the honest cost metric.