JournalAgentic AI
Agentic AI and the Compounding Cost of a Wrong Turn
Per-step accuracy is a seductive metric. At 95% per step, a twenty-step task succeeds only about a third of the time.
The pitch for agentic systems is that the model stops answering and starts doing. The engineering reality is that autonomy converts a single prediction problem into a sequential one, and sequential problems multiply their errors.
The arithmetic nobody puts on the slide
steps p=0.95 p=0.99
5 77.4% 95.1%
10 59.9% 90.4%
20 35.8% 81.8%
50 7.7% 60.5%A 95% step is an excellent model and a poor agent. This is why demos that chain five tool calls feel magical and production workflows that chain thirty feel broken — both are behaving exactly as the arithmetic predicts.
Three levers that actually move the number
- Shorten the chain. Every step you can replace with a deterministic function is a step that cannot fail probabilistically. Most agent traces contain more of these than teams expect.
- Make steps verifiable. A step whose output can be checked cheaply — schema validation, a compile, a test run, a database constraint — converts a silent failure into a retry.
- Checkpoint state. If step 18 fails, resuming from 17 costs one step; restarting costs eighteen and burns the user's patience.
Autonomy is not a feature you add. It is a budget you spend, and the currency is verifiability.
What to instrument
Aggregate task success hides everything you need. Log per-step outcomes and you can see which step is eating the budget:
- Step-level success rate, not just end-to-end completion.
- Retry depth — a step that usually succeeds on attempt three is a step that is failing.
- Divergence: how far the executed plan drifted from the plan proposed at step zero.
- Human interventions per completed task. This is the honest cost metric.