Back to Resources
AI Development Engineering 7 min read

The One-Model Trap: Why Your LLM Bill Is Higher Than It Should Be

Most startup engineering teams route every AI request to their most capable, most expensive model and call it a decision. This post examines what that default actually costs, using real pricing data from 2026's model landscape, and makes the case for treating model routing as an infrastructure concern from the start. Readers walk away with a clear framework for categorising their own traffic and a practical guide to tooling that fits where they actually are

15 Sep 2026

The invoice came on a Tuesday. $8,400 for 31 days of inference. The founder forwarded it to his lead engineer with one word: “Explain.” The engineer pulled up the logs and found it: 17,000 times in one month, their most expensive model had been asked to extract a job title from an email signature. A task requiring about as much intelligence as a dictionary lookup. Billed at frontier rates.

This is not a horror story about runaway AI. It is a story about a default that nobody made consciously.

The Architectural Decision You Didn’t Know You Made

When a team builds its first AI feature, the model selection process goes roughly like this: the engineer tries a few options, picks the one that gives the best results on the tasks they care about, and hardcodes it. That is a reasonable approach for a prototype. The problem is that prototypes have a way of becoming production.

Six months later, that hardcoded model is handling everything: complex reasoning, document classification, text extraction, tone detection, and yes, pulling a job title out of an email signature. It is the same right-sizing problem we cover in You Don’t Need to Fine-Tune That — matching the tool to the task instead of defaulting to the most expensive option available. The engineer who made the original choice has moved on to three other problems. Nobody looks at the line item closely until the bill doubles.

This is what I call the one-model trap. Not a mistake exactly. A default. A decision made by omission, at a time when the right answer was “good enough” and the wrong answer was invisible.

By 2026, more than 75% of engineering teams are running multiple models in production environments. But most of them arrived there reactively, after the bill arrived, not before. The gap between “we have a thoughtful multi-model architecture” and “we panicked and added a cheaper fallback last quarter” is where most startups live.

LLM Model Routing Cost Comparison: Single-model vs routed architecture showing cost savings by request type

The Math That Makes This Uncomfortable

The price spread between models in 2026 is roughly 100x. At the low end, smaller specialist models run at around $0.09 to $0.44 per million input tokens. At the high end, frontier reasoning models run at $15 to $30 per million input tokens, with output priced higher still.

Here is a rough example. Suppose your system processes 50 million input tokens and 10 million output tokens per month. If everything goes through your most capable model at $5 input and $25 output, the monthly bill comes to around $500.

Now suppose you route 80% of that traffic, the commodity requests, to a smaller model at $0.09 input and $0.18 output. The remaining 20% still goes to the frontier model. The new monthly total: roughly $55.

An 89% reduction. Without changing a single user-facing feature.

The teams reporting 40 to 85% cost reductions from model routing are not doing anything exotic. They are just making a distinction that most teams skip: not every request needs your smartest model.

What Routing Actually Is

Model routing is the decision layer that sits between your application and your model provider APIs. Your application makes a single call. The router decides which model handles it, based on rules you define.

Those rules can be simple: anything under 200 tokens goes to the cheap model, anything above goes to the expensive one. Or they can be more sophisticated: classification requests go to model A, summarisation goes to model B, and anything flagged as high-stakes goes to model C. The router can also handle fallbacks, sending a request to a secondary model if the primary times out or errors.

There are three common routing patterns:

Resilience routing keeps the system running when a provider has an incident. The router detects a failure and falls back to an alternative. This is the minimum viable version of routing, and honestly, more teams should have it regardless of cost.

Cost routing reduces spend by matching each request to the cheapest model that can handle it at acceptable quality. This is where the real savings live.

Quality routing goes further: it measures actual output quality from production traffic and uses those scores to decide which model gets which requests. This is the most sophisticated version, and for most startups, it is also overkill until you have enough volume to make the measurement meaningful.

For a seed-to-Series-B team, cost routing with resilience fallbacks is the practical target. Quality-based routing is a nice problem to have later.

The Three Types of Traffic You Probably Have

Most production AI systems, once you actually look at the logs, have a recognisable shape.

Commodity tasks are high-volume and low-complexity. Classification, entity extraction, intent detection, simple transformation, format validation. These tasks do not require a frontier model. A well-prompted smaller model handles them at a fraction of the cost, and the output quality difference is negligible at scale.

Standard tasks are your workhorse use cases: summarisation, structured output generation, moderate reasoning, customer-facing text that needs to be good but not exceptional. A mid-tier model does this well. Think Gemini Flash, Claude Haiku, or the smaller versions of whatever your preferred provider offers.

Precision tasks are the ones where quality actually matters and errors are expensive: complex reasoning, high-stakes decisions, generation that a human will rely on without review, anything where a wrong answer has downstream consequences. These deserve your best model. But they are probably a smaller slice of your traffic than you think.

When teams audit their actual request distribution, they typically find that 60 to 80% of their traffic falls into the commodity category. That is the traffic being sent to a frontier model and billed at frontier rates.

When Routing Makes Sense (and When It Does Not)

Routing adds engineering overhead. There is a router to maintain, routing rules to define, and fallback behaviour to test. It is not free.

The conventional wisdom is that routing pays for itself above roughly 100,000 daily active users, or a monthly inference bill around $10,000. Below that threshold, the engineering cost of building and maintaining a routing layer can outpace the savings.

That said, I would argue the threshold question is slightly wrong. The better question is: are you building an architecture you will have to rebuild anyway in 12 months? Even if your current traffic does not justify routing today, designing your AI layer with routing in mind costs very little at build time. Adding it after the fact, when you are already on a production system handling real users, costs significantly more.

The teams I have seen handle this well treat model selection as an infrastructure concern from the start, even if the initial implementation is simple. One model, one fallback, clear rules. That is enough to avoid the locked-in-to-one-provider problem and to set the team up to route more intelligently later.

Our AI consulting work often starts here: not with “which model should we use?” but with “how should your AI layer be structured so you can change your answer to that question without rebuilding everything?”

Picking Your Tools

The routing tooling in 2026 is genuinely good, and most of it is either open-source or has a generous free tier.

LiteLLM is the open-source standard for teams that want to own the routing layer. It provides an OpenAI-compatible proxy that routes across 100+ models, with support for virtual keys, budgets, rate limiting, and fallbacks. The engineering lift to self-host it is real but manageable for a team with platform experience.

OpenRouter is the fastest path to broad model access with minimal setup. 400+ models across 60+ providers via a single API. It handles provider-level routing and fallbacks but does not include quality-based routing natively. Good starting point if you want to experiment before building something custom.

Portkey and Vercel AI Gateway both sit in the managed middle ground: more configuration control than OpenRouter, less infrastructure ownership than LiteLLM. Both handle conditional routing, caching, and spend visibility.

Braintrust is the option to reach for when you want routing decisions backed by actual quality measurement from production traffic. It connects gateway routing to evaluation scores and tracing, so you can see not just which model is cheapest but which model is cheapest at a given quality bar. More setup, more signal.

The choice depends less on features and more on what your team can maintain. An open-source proxy you cannot keep up to date is worse than a managed service with fewer options. Our AI-powered development work includes routing layer design as a standard part of production architecture, because the tooling decision and the architecture decision are inseparable.

The Longer Risk: Vendor Lock

Cost is the visible reason to implement routing. The less visible reason is resilience.

When your system routes exclusively through a single provider, that provider’s pricing decisions, API changes, and outages become your system’s problems. OpenAI changed its API structure three times in 2024. Providers have deprecated model versions with 90-day notice windows. If your entire inference stack depends on one endpoint, any of those changes requires a coordinated engineering response under time pressure.

A routing layer that sits between your application code and your providers is the architectural equivalent of a seam. When the provider changes something, you update the routing configuration, not the application. When a provider has an outage, you reroute traffic, not rebuild.

This is the same argument that makes workflow automation valuable as a structural approach: building the flexibility to swap components without rewriting the thing that depends on them. A routing layer is that pattern applied to model selection.

The teams doing this well are treating LLM providers the way they treat cloud regions: you commit to an interface, not a destination. Your application does not know or care which model handled a request. The router decides, based on current conditions, cost, and quality signals.

The Honest Version

None of this is magic. Routing introduces complexity. Routing rules can be wrong. A cheaper model that handles 80% of your commodity traffic well still fails on the 20% you misclassified. Quality monitoring on routed systems is harder than monitoring a single model. These are real costs.

But the alternative is not a simpler system. It is a simpler-looking system with a hidden cost structure that will surface either in your cloud bill or in a provider dependency you discover at the worst possible moment.

The invoice that came on that Tuesday could have been $55. The engineer who dug through the logs knew it by the time he sent his reply. His next sprint was a routing layer.

Knowing what to build before the bill arrives, and building it in a way that lets you change your mind later: that is what production-ready AI architecture actually looks like.

Previous article The Bill Arrives in Year Two AI Development Next article Passing Tests Is the Easy Part AI Development
Related articles
AI Development Software Engineering 6 min read

The Accuracy Tax: What Chaining AI Steps Actually Does to Your Pipeline

The demo looks flawless. A five-step agentic workflow: pull the contract, extract the key clauses, flag the risk terms, draft...

AI AI Agents Software Development
Read
AI Development Engineering 6 min read

Passing Tests Is the Easy Part

75% of AI coding models introduce regressions when maintaining real codebases over time. Here's what the SWE-CI benchmark reveals and what your engineering team should do about it.

AI AI Evaluation Software Development
Read
AI Development Engineering 6 min read

The Bill Arrives in Year Two

AI-generated code now makes up 41% of new code. The teams that moved fastest in 2024 are hitting a year-two maintenance wall. Here is what the data shows and what the teams managing it well actually do.

AI Production Ready AI Software Development
Read