The product demo was going well until the form.
A VP of Engineering at a mid-stage SaaS company was showing off their new AI-powered customer onboarding flow. The assistant asks a few questions, extracts the user’s company name and billing preferences, and fills in a summary. Clean. Useful. Exactly the kind of thing investors like to see.
Then he clicked “Submit” and they waited.
Five seconds. Ten. The spinners kept spinning. At the seventeen-second mark, the AI finished “thinking” and populated three fields. The room stayed quiet. One of the investors said, “Is it always that fast?”
The assistant had been routed to a reasoning model. An o4-class model doing extended inference on the task of parsing a company name and a billing cycle from two text fields. It produced the right answer. It just spent the cognitive equivalent of solving a chess endgame to do it.
This is the thinking tax. And right now, a lot of engineering teams are paying it without realising why their infrastructure bill climbed and their users stopped finishing onboarding.
What Reasoning Models Actually Do
Reasoning models are genuinely different from the generation of models that came before them. Where a standard model reads your prompt and generates a response in a single pass, reasoning models run extended internal chains of thought before arriving at an answer. They deliberate. They check their work. They try alternate paths and abandon ones that don’t hold up.
The results on hard problems are real. o3 scored 87.7% on graduate-level science questions. These models catch subtle logic errors, handle multi-step proofs, and can navigate complex code refactors across dozens of files in ways that their predecessors could not.
But that deliberation is not free. A single hard problem can consume tens of millions of internal tokens before producing a final answer. Response times stretch from the sub-second range of standard models into fifteen to forty-five seconds. And the cost difference at production scale is significant enough to show up in your AWS bill alongside the things you thought were expensive.
The question is not whether reasoning models are impressive. They are. The question is whether your specific task actually requires reasoning.

The 85 Percent Problem
Here is a rough breakdown of what most production AI applications actually spend their time doing: classifying inputs, extracting structured data from text, summarising documents, generating copy, answering conversational questions from a knowledge base, and doing simple question-answer lookups.
Standard models handle 85 to 92 percent of real-world tasks with 200 to 500ms latency and predictable per-token costs. When you route those tasks through a reasoning model instead, you get the same answer, sometimes worse copy because the model overthinks phrasing, and you pay three to five times more for the privilege.
A production system handling 10,000 daily requests costs $30 to $100 per day on a well-routed stack. Run those same requests through reasoning models and you are looking at $150 to $300 per day. That gap does not feel painful when you are testing in staging. It becomes painful when you are three months into Series B runway and your inference costs are climbing faster than revenue.
This is the scenario we see most often when working with engineering teams on AI-powered development. The model selection choice gets made during a proof of concept, where cost does not matter. Then the feature ships, usage scales, and nobody goes back to ask whether a faster, cheaper model would have served just as well.
Where Reasoning Actually Earns Its Cost
To be clear, there are situations where routing to a reasoning model is the right call, and trying to save money by routing those to a standard model will cost you in other ways.
Complex multi-file code refactors are one. When you are asking an AI to understand a sprawling codebase, identify all the places a pattern needs to change, and do that correctly without breaking interfaces, the extended thinking is doing real work. Standard models produce plausible-looking but subtly wrong refactors at a rate that makes them more expensive in practice once you factor in the engineering time to find and fix errors.
Novel problem-solving is another. If you are building an AI system that needs to figure out how to handle a case it has not seen before, or plan a sequence of tool calls across a complex workflow, the deliberation pays off. Reasoning models are better at recovering gracefully when their first approach hits a dead end.
High-stakes analysis with complex tradeoffs also fits, especially when the output will be reviewed by a human anyway. A legal clause review, a risk assessment, a technical architecture evaluation. In these cases, the response time is not a real constraint, and the extra accuracy is worth having.
The pattern that scales well: route 80 to 90 percent of your traffic to a standard model and reserve the reasoning-heavy tier for the 10 to 20 percent of requests where it genuinely matters. This is essentially the same logic covered in our earlier piece on the one-model trap, applied specifically to inference-time compute.
The Auditability Problem Nobody Talks About
There is a second cost that does not show up in your invoice.
Reasoning models produce opaque decision paths. When the model thinks for 45 seconds internally and then gives you an answer, you have the answer but you do not have a clear account of how it got there. For most consumer applications, this is fine. For FinTech, HealthTech, and any domain where you need to explain a decision to a regulator, it is a serious problem.
A standard model with a well-structured prompt produces a decision that can be audited. You can show the prompt, show the output, explain the reasoning in natural language, and a compliance review can follow the chain. That auditability is part of what you are paying for in those contexts.
If you are building workflow automation in a regulated space and you default to reasoning models because they score better on benchmarks, you may find yourself unable to explain a decision during an audit, not because the decision was wrong, but because the model’s internal reasoning chain is not surfaced in any way that maps to what compliance reviewers need.
This is a design choice, not a model limitation. But it is one worth making deliberately rather than discovering after a compliance review.
The Question Engineers Should Ask First
Before you reach for a reasoning model, ask one question: is this task failure domain hard, or is it execution domain hard?
Failure domain hard means the task is conceptually difficult. The logic is deep, the problem is novel, the path to the right answer is not obvious. These are the tasks where reasoning earns its cost.
Execution domain hard means the task is technically complex to wire up, but the underlying reasoning is straightforward. Extracting fifteen fields from a PDF is technically involved but conceptually simple. Classifying a support ticket into one of forty categories requires robust engineering around the model but not a reasoning engine.
Most teams, when they look honestly at their task mix, find that the vast majority of what their AI does is execution domain hard. The reasoning model is solving for the wrong thing.
Practical First Steps
If you are not sure whether your current model routing is costing you unnecessarily, the diagnosis is not complicated. Run a sample of 200 to 500 representative production requests through both a standard model and your current reasoning model. Compare the outputs on a quality rubric appropriate to your task. If the standard model is within acceptable error bounds at a fraction of the cost and latency, you have your answer.
If you are starting a new AI feature, the default should be a standard model with a well-designed prompt. Upgrade to a reasoning model only for the specific tasks where you can demonstrate the quality lift is worth the cost. This is a core part of how we approach AI consulting for engineering teams who are trying to build things that stay economical at production scale.
Reasoning models are a real capability gain for the problems they were designed to solve. The issue is not the models. It is the reflex of reaching for the most powerful tool in the stack regardless of whether the problem calls for it.
The thinking tax is optional. Most teams just do not realise they are paying it.