Category Archives: Cost

Stop Reaching for the Biggest Model

Stop Reaching for the Biggest Model

There is a reflex that shows up in nearly every generative AI design review I sit in. Someone needs to classify support tickets, or extract fields from invoices, or summarize a document, and the default choice is whatever frontier model is topping the leaderboard this quarter — Claude Opus, GPT-5.x, Gemini. It is the obvious pick. It is the safe pick. And it is frequently the wrong pick.

The problem is not that frontier models are bad. They are extraordinary. The problem is that you are about to multiply that per call premium by millions of calls, and nobody in the room has asked whether the job actually requires a frontier model at all. The right question is almost never “what is the best model?” It is “what is the cheapest architecture that clears my quality threshold reliably?”

Those are very different questions, and they lead to very different bills.

The economics nobody runs the numbers on

Let me make this concrete with a workload I see constantly: summarization at scale. Suppose you run 10,000 summaries a day. Each request is roughly 1,000 input tokens and 300 output tokens — a realistic shape for document or thread summarization.

Run that on a small, task specific model in the 2B to 14B parameter range and you are looking at something like $7.20 a day. Run the identical workload on a mid tier frontier option like GPT-4.1 Mini and you land closer to $58 a day. Same inputs, same outputs, and — this is the part that matters — in the benchmark those numbers come from, human evaluators could not detect a quality difference in the summaries themselves.

That gap is about $18,500 a year on a single workload. Not your whole platform. One job. Most organizations have a dozen of these running quietly in production, each one silently overpaying because the model choice was made once, early, by reflex, and never revisited.

The AWS Machine Learning blog made this point sharply in Beyond the price per token: the sticker price on a token is not the number that matters. What matters is the cost of a completed, acceptable task. And when the task is narrow and measurable, a smaller model routinely wins that comparison outright — lower cost, lower latency, and quality that is indistinguishable from the expensive option.

The honest counterexample

Now, if I stopped here I would be selling you a lie by omission. Small does not mean smart. Small means small.

Consider reasoning. On GSM8K — a grade school math word problem benchmark that demands multi step reasoning — a 0.5B parameter model scored around 37.7%. Llama-3.1-70B scored 92.0% on the same benchmark. That is not a rounding error. That is the difference between a system you can ship and a system that is wrong on nearly two thirds of its answers.

So this is not “small always wins.” This is task matching. The summarization job and the math reasoning job look superficially similar — both are “send text, get text” — but they place completely different demands on the model. Summarization is largely a compression and rephrasing task that small models handle beautifully. Multi step arithmetic reasoning is exactly the capability that scales with parameter count, and starving it produces garbage.

The discipline is knowing which is which. High volume, narrow, well defined work — classification, extraction, routing, summarization, format conversion — is small model territory. Open ended reasoning, complex code generation, nuanced judgment across long context — that is where you pay for the frontier, and you should. Both Refonte and Forbes have documented small language models beating frontier systems on cost, speed, and accuracy — but always on the tasks that fit them.

The discipline: measure completed task economics

If you take one thing from this post, take this: stop comparing API price cards. Price per token is an input, not an outcome. What you actually pay for is a completed task that meets your bar, and that number includes things the pricing page never shows you:

  • Accuracy at your threshold. Not headline benchmark scores — accuracy on your data, at your acceptable quality line.
  • Retries. A cheaper model that fails and gets rerun twice is not cheaper.
  • Human review rate. If a smaller model pushes 15% of outputs to a human reviewer and a larger one pushes 3%, the reviewer’s time may dwarf any token savings.
  • Escalation cost. What does a wrong answer cost downstream when it slips through?

When you sum those up, the winning model is often not the one with the lowest token price or the highest benchmark score. It is the one that clears your threshold with the least total cost per acceptable output.

The AWS Well-Architected Generative AI Lens codifies this as GENCOST01-BP01: Right-size model selection. The guidance is refreshingly blunt — start with the smallest model that could plausibly work, and scale up only when your evaluations show you need to. This is the inverse of the reflex. Instead of starting at the top and hoping you can justify the cost, you start at the bottom and earn your way up with evidence.

A pattern that makes this operational is complexity based routing. Rather than sending every request to one model, you inspect the request and route it: simple ones to a small fast model, hard ones to the frontier. On AWS you can do this with Amazon Bedrock Intelligent Prompt Routing, which dynamically dispatches each prompt to the most cost effective model that can handle it.

incoming request
      |
      v
  classify complexity  ──> simple / narrow  ──> small task specific model (2B–14B)
      |                                              |
      |                                              v
      |                                        meets threshold? ── yes ──> done
      |                                              | no
      v                                              v
  complex / open-ended ──────────────────> frontier model

And critically: right sizing is not a one time decision. New models ship every few weeks, and the price to performance frontier moves under you constantly. Treat model selection as a continuously revisited parameter, not a setting you configure once and forget.

Practical takeaways

  1. Ask the right question. Not “what is the best model?” but “what is the cheapest architecture that clears my quality threshold reliably?”
  2. Measure completed task economics, not price per token. Fold in retries, human review rate, and escalation cost.
  3. Start small and scale up. Follow GENCOST01-BP01: begin with the smallest plausible model and only move up when evaluations demand it.
  4. Route by complexity. Use Bedrock Intelligent Prompt Routing so simple requests get small model pricing and hard ones still get frontier quality.
  5. Re-evaluate continuously. The efficient frontier moves. Revisit your model choices as new options ship.

Paying for general purpose capability you never use is not a convenience. It is an architectural choice — and for high volume, narrow, measurable work, it is usually the wrong default…

Stop Reaching for the Biggest Model

There is a reflex that shows up in nearly every generative AI design review I sit in. Someone needs to classify support tickets, or extract fields from invoices, or summarize a document, and the default choice is whatever frontier model is topping the...