The demo looks flawless. A five-step agentic workflow: pull the contract, extract the key clauses, flag the risk terms, draft the summary, route it to the right reviewer. Clean inputs in a staging environment, and it runs perfectly four times out of four. The product manager is delighted. The investor asks if you can do ten steps instead of five. You say you will look into it.
Then you ship it.
Three weeks in, the workflow is misfiring on roughly one in three real contracts. Not catastrophically. Nothing explodes. The summaries are slightly off, the risk flags occasionally miss a clause that a junior lawyer would have caught, and the routing sends documents to the wrong team twice in the first week. Nobody can pinpoint which step is breaking. Every step, tested in isolation, looks fine.
This is not a bug. It is arithmetic.
The Math Your Demo Was Hiding
Here is a number most teams never calculate before they build: if each step in a pipeline succeeds with 90% accuracy, a five-step pipeline succeeds end-to-end about 59% of the time. At ten steps, that same 90% per-step accuracy delivers a correct result roughly 35% of the time. You have built a system that is wrong more often than it is right, out of components that each look excellent in isolation.
The formula is p^n. If each step succeeds with probability p, then n sequential steps succeed at p raised to the power of n. At 85% per-step accuracy across ten steps, you are at 20%. At 80%, you are at 11%.
This is not a critique of AI quality. A 90% success rate on a discrete task is genuinely good. The problem is that pipelines multiply errors rather than absorb them.
Teams understand this intuitively with human processes. A document that passes through five reviewers, each of whom catches 90% of errors, leaves ten times the residual error of a single expert review. No one is surprised by that. But when AI replaces each reviewer in the chain, something about the automation makes the same math feel like a detail. It is not. It is the entire design problem.

The Second Problem Is Worse
Multiplying step failures is manageable if errors are independent. Often they are not.
Research from 2025 on multi-agent reasoning systems documented what the authors called “self-conditioning degradation.” When a model’s context window contains its own earlier mistakes, it becomes measurably more likely to produce further mistakes on downstream steps. The error does not simply stay contained. It actively degrades the reasoning context for every step that follows. And that degradation accelerates as the pipeline grows longer.
Think about what this means concretely. Step three of your pipeline extracts a clause incorrectly. Step four, which drafts the risk summary, inherits that bad extraction and builds on it. Step five, which handles routing, is now working from the compounded confusion of steps three and four together. You are not looking at three independent failure modes. You are looking at a single error that has been amplified twice before it reaches the end.
At ten steps, that dynamic can take a 90% accurate first step and produce an output that fails more than half the time, because each prior failure makes the next failure more likely. The math compounds. So does the damage.
What This Looks Like Inside a Real Team
The failure pattern is recognizable, once you have seen it. You build an agentic workflow where each piece tests solid. QA reviews every step in isolation. The pipeline ships to production. Then the edge cases arrive: unusual document formats, ambiguous inputs, user behavior that staging never surfaced. And instead of a single predictable failure mode, you get a fog of partial, cascading errors that resist tracing back to any one origin.
Engineers spend weeks adding logging. The logs show everything running at 88-92% accuracy per step. Nobody finds the bug because there is no single bug. The system is doing exactly what it was designed to do. The design just did not account for what multiplication does to accumulated error at scale.
This pattern surfaced in a 2025 study of multi-agent coordination systems. Adding relay agents without introducing genuinely new information degraded accuracy from 90.7% to 22.5%, below what a single well-prompted model would have delivered on its own. The extra agents were not helping. They were compounding each other’s uncertainty.
The same study found that decentralized peer architectures, multiple agents consulting each other in a mesh, amplified errors 17.2 times compared to a single-agent baseline. A centralized architecture with one coordinator contained the same error multiplication to 4.4 times. Neither number is zero, but the difference between them is the difference between a system you can operate and one you cannot.
Topology matters more than most teams realize when they are planning the pipeline. This is exactly the kind of architectural conversation that should happen during AI consulting before a single line of code is written, not during an incident review after three weeks of production data.
How Teams Get Pulled Into the Trap
The incentive structure pushes toward longer pipelines. A three-step process sounds less impressive than a seven-step one in a pitch deck. Each step you add creates an opportunity to claim a new capability. And when the staging environment is cooperating, the pipeline does look better at each step you add.
What staging environments rarely simulate well: the diversity of real inputs, the edge cases that users generate naturally, the slightly-malformed data that is common at any scale, and the accumulated drift that happens when a model is running in production for weeks instead of hours. The test set that your team assembled to validate the pipeline is, by definition, missing the things that will cause real failures.
There is a version of AI-powered development that rushes to add steps because each step looks like a feature. And there is a version that pauses after the first step and asks: what is the real-world accuracy on production-shaped inputs, before we build anything on top of this? The second version ships slower and breaks less.
What Actually Works
The agentic patterns that held up in production through 2026 share two characteristics: they are linear with clear intermediate artifacts, and they have checkpoints where a confident gate can catch compounding drift before it propagates further.
The patterns that failed are the ones that look more impressive on a whiteboard. Open-mesh agents collaborating freely. Long chains with no intermediate validation. Distributed architectures where every step trusts every other step’s output implicitly.
Concrete guidance that applies to almost every team building this right now:
Start with one well-prompted model doing the entire task end-to-end. Measure its real accuracy on a representative sample of production-shaped inputs, not a curated test set. If you hit a genuine bottleneck because the task requires distinct reasoning in parallel domains, add one more agent for the clearly separate domain. Not because it will look more sophisticated. Because the task shape requires it.
When you do chain steps, build hard gates. A step that produces output below a confidence threshold should halt the pipeline and route to human review, not silently pass degraded output to the next stage. Soft failures in AI pipelines become hard failures nine steps later.
Keep the chain short. A five-step pipeline is not twice as capable as a three-step one. Most tasks that teams decompose into ten steps can be done better with three steps and better prompting on each one. The complexity budget should go into making each step more reliable, not into adding steps that multiply the existing uncertainty.
The Measurement Problem You Need to Solve First
There is a reason teams discover the accuracy tax late: they are measuring the wrong thing. Most pipelines get evaluated at the final output. Either the document was routed correctly or it was not. Either the summary was accepted or flagged for revision. That final-output view hides which step introduced the problem, and it hides how close you are to the floor.
What you actually need is accuracy measurement at every intermediate step, with a sample of real production inputs rather than synthetic ones. That means logging the output of step two alongside a ground-truth label, even when step two’s output is not user-facing. It means setting up a lightweight eval loop that runs a subset of live inputs through each step in isolation, weekly, and flags drift before it shows up as a customer complaint.
This is the part that slows teams down. Building the pipeline is fast and satisfying. Building the instrumentation to know when the pipeline is degrading is slower and less glamorous. Most teams skip it until they need it. At that point, they are debugging a fog instead of a number.
Before You Add Another Agent
The practical test before any pipeline expansion: if you run this new step one thousand times on real production inputs, what is the accuracy? Not on the examples you hand-picked for the demo. On the messy, ambiguous, occasionally malformed inputs that users actually produce.
Multiply that accuracy rate by the accuracy of every upstream step. That number is your real end-to-end success rate, before you add anything. If it is below 80%, you do not have a pipeline problem that more agents will fix. You have a step problem, and adding more steps will deepen it.
The teams that build agentic workflows that hold up are the ones who treat the pipeline as a system with compounding failure modes, not as a sequence of individual features that each happen to work. They measure accuracy at every gate. They build halt conditions before they build the happy path. They resist the pull to add a sixth step because the first five look clean in testing.
The math is not forgiving. But it is entirely predictable. And a predictable problem, once you can see it, is one you can actually design around. Build the pipeline your inputs deserve, not the pipeline your demo made possible.