Choosing a Claude model in Claude Code: why the cheaper one is usually the expensive one

Picking a model in Claude Code looks like a price decision and isn't. It is two independent dials — which model, and how hard it thinks — and the only number that should move either is the cost of finishing the whole task, not the price on the token. This report explains why a cheaper model is so often the expensive one (the turn-count mechanism that makes weak models superlinearly costly on hard work), how to set both dials from four properties of the task, why the effort dial pays off on different curves for different work, and why a benchmark score is a weak guide when real performance is model × harness × product.

Published · By Dan Walter

Executive summary

In Claude Code, the cheaper model is usually the more expensive one. That reads like a contradiction only because the wrong number is doing the deciding. Choosing a model is not picking a line off a price list; it is setting two independent dials — which model, and how hard it thinks — and the number that should move both is the cost of finishing the whole task, correctly, once. A weaker model turned loose on ambiguous work does not do less work; it does more — more exploration, more retries, more circling — and because every step of an agentic task resends the whole conversation so far, the cost of a task climbs with something closer to the square of the turns it takes. The low sticker price quietly buys the high bill.

From that one correction the rest follows. The model dial sets the intelligence ceiling; the effort dial sets how much of that ceiling gets spent on this particular job. They answer different questions and move independently, and the right setting for each is read off the task — its ambiguity, its horizon, how easily a mistake is caught, and how often the job repeats — starting from a capable default, stepping down only where quality visibly holds and up the moment it slips. Anthropic's own measurements point the same way: on work that genuinely stretches a smaller model, the more capable model finishes both better and cheaper per task. The effort dial, meanwhile, pays off on a steep curve for building and long-horizon work but a nearly flat one for research and chat, so the identical setting is a bargain in one place and pure waste in another.

The last correction is the one that settles most arguments. A model's real-world performance is not a property of the model. It is model × harness × product: the same model backend swings roughly twenty-four points of task success depending only on the software wrapped around it. A leaderboard number, measured in someone else's harness on someone else's tasks, is therefore a shortlist and never a verdict — and everything here is written about Claude Code specifically, which is itself a harness. The whole discipline reduces to a single habit: stop optimising the price of the token and start optimising the cost of the finished job, judged on the surface where the work actually runs.

You are not choosing a model — you are setting two dials

Nearly everyone treats this as one knob labelled model, usually with a pricing page open and a worry that the expensive option is a waste. There are two knobs, they move independently, and the confusion between them is the root of most of the money spent badly. Getting them right matters more the more of a working day runs through these tools: for an operator who lets agents do the volume work, a mis-set dial is not a one-off overcharge but a tax levied on every task, all day.

The first dial is the model. From most to least capable: Fable, Opus, Sonnet, Haiku. Anthropic's own shorthand is telling — Fable is "a specialist who's seen problems almost no one else has," Opus is "the expert," and Sonnet is "a really good generalist." Capability descends the list and so does price, but the dial is really setting one thing: the intelligence ceiling for the work. Larger models are also markedly better at handling ambiguity, while smaller ones do their best work when the instructions are specific and the path is spelled out — a property that turns out to decide most of the choice.

The second dial is reasoning effort — a separate control, running low, medium, high, xhigh, max, that governs how hard the model thinks and how much it explores before it acts. If the model sets what is possible, effort sets how much of the possible gets paid for this time. The two come apart in practice: a top model can be run lazily and a small model flat-out, and those are entirely different choices with entirely different bills. There is a tell worth noticing — the effort dial only exists on the frontier models. On the smallest one it is not offered at all, which is itself a clue about what that model is for.

Because the dials are independent, they multiply into four very different choices: a top model run lazily, a top model run flat out, a small model barely thinking, a small model working hard — each with its own bill and its own failure mode. Claude Code exposes them as two separate controls, set with their own commands, and effort in particular is better treated as a standing preference for the kind of work at hand than as a knob fiddled with per task. The common mistake is not picking the wrong value; it is collapsing the two dials into one and imagining a single slider that runs from "cheap" to "good." The clearest way to hold both at once is to lay the work out across them.

Exhibit — Most work sits low and left; only the hardest, most open-ended jobs earn the top-right corner.. Two dials, not one — and 'maximum everything' is the right home for very little of the work. Source: Map of the two dials in Claude Code, per Anthropic, 'Choosing a Claude model and effort level in Claude Code' (2026). Positions illustrative.

The map is the argument: two dials, and a great deal of empty space in the expensive corner. Which raises the only question that matters — if not price, what decides where a job belongs?

The number that decides is the cost of the finished task, not the price of the token

Here is the reframe, and it is the whole game. The price per token is a rate. It says nothing about how far the model gets before the job is done. The number that actually matters is the cost of the finished task — the whole thing completed, correctly, once — and those two numbers come apart far more than intuition expects, in the direction that punishes the "cheap" choice.

A weaker model on an ambiguous or complicated job does not do less work; it does more. More back-and-forth, more retries, more tokens spent going in circles, and more human time cleaning up afterwards — and human review is the single most expensive destination in the whole system. A downgrade, in other words, does not reduce the work. It reprices it, upward, because the model was the cheapest worker in the loop and the alternative to its next turn is a person.

The mechanism underneath is not linear, which is exactly why it surprises people. Every step of an agentic task resends the entire conversation so far, so a job that runs forty turns sends its opening context something like forty times over. The cost of a task therefore grows with roughly the square of the number of turns it takes. A cheaper model that flails and needs twice the turns is not twice as expensive — it is worse than that. This is the real reason a smaller model "becomes expensive" the instant ambiguity appears: not the per-token price, the turn count. The same effect runs through Claude Code's other big cost drivers, which are all quantity-of-context effects — the files read into the window, the length of the session, the fan-out to sub-agents — not the unit price of the model at all.

This is not a hunch. Anthropic's own guidance states it plainly: on tasks that genuinely stretch the smaller model, "the total cost per task can come out lower" on the larger one despite its higher per-token price. Their measured comparisons show the same shape — the more capable model finishing hard problems both better and cheaper per task — and independent analyses reproduce it from the outside, finding a lower-priced model completing a benchmark at a higher cost per answer because it burned more output and took several times the turns. The cheaper token produced the bigger invoice. Plot the tiers on the two things that actually matter and the shape gives the game away.

Exhibit — On complex work, the cheapest token buys the dearest finished job.. Chasing the cheapest token walks you away from the corner you want — the one that finishes the job for least. Source: Directional shape of Anthropic's published cost-per-task measurements on hard tasks ('Choosing a Claude model and effort level in Claude Code', 2026), not exact figures.

Notice the corner everyone aims for — cheap by both measures — sits empty on hard work. The lowest finished-task cost is in the middle of the price range, not at the bottom of it.

There is a subtler trap in averages. The bill on a batch of agentic work does not sit in the median task; it sits in the tail. In one of Anthropic's research runs, two problems out of twenty carried more than forty percent of the total spend — the handful that spiralled into long, uncertain loops. A cheaper model tends to look best exactly where being wrong is cheapest, on the easy median task, and to lose precisely where the money is, on the hard tail it turns into a marathon. Benchmark a tier on its average task and the comparison is run on the cases that were never going to cost anything; the honest test is what the worst few tasks in a real workload cost end to end.

So the choice is never "what is the cheapest model?" It is "what finishes this task, correctly, for the least total effort?" Those are different questions, and mistaking one for the other is the single most expensive error in the whole practice.

Choose both dials from the task in front of you, not the price list

Once the right number is doing the deciding, the rule is simple, and it is the one both Anthropic and OpenAI quietly converge on: optimise for getting the job done first and for cost second. Start from a capable default for the kind of work it is. Step down a rung only after watching the quality hold, and step up the instant the output goes shallow, inconsistent, or wrong on the cases that matter.

Stepping down is not a leap of faith; it is cheap to test. In Anthropic's coding comparisons the second-tier model came within a fraction of a point of the top one at roughly sixty percent of the cost — when a downgrade holds like that, take it. When it does not hold, the failures announce themselves on exactly the cases that would have been caught in review anyway. And a handful of properties of the task tell you which way to lean before a single token is spent.

Exhibit — Four signals tell you which way to lean — up on ambiguity and horizon, down on checkable, high-volume work.. Start capable, step down only where quality holds, and step up the moment it slips. Source: Framework synthesising Anthropic ('Choosing a Claude model and effort level in Claude Code', 2026) and OpenAI model-selection guidance.

The four signals point the same way for a reason: each is really asking how likely is this to spiral into a long, uncertain loop? The higher that risk, the more a capable model pays for itself by reaching the end in fewer turns — the turn-count mechanism, run in reverse. There is a productive tension in the advice worth naming: some guides say start high and step down, others say start cheap and step up. They are the same principle seen from two ends, and which direction to walk depends on which unknown is bigger — whether the task is even feasible at the cheaper tier, or merely cheaper. When feasibility is in doubt, start high and prove you can come down. When only cost is in doubt, start low and escalate on the first bad case. Either way the task, not the price, sets the opening bid.

Of the four signals, ambiguity is the master, because it is the one the models themselves are most unequal on. The larger models are distinctly better at holding an under-specified problem and finding the intent inside it; the smaller ones do their best work when the goal is spelled out and the path is closed, which is why a precise, step-by-step prompt rescues a small model and a vague one exposes it. That single fact resolves most of the choice — a wide, exploratory brief wants a capable model almost regardless of the other three signals, and a crisp, mechanical one can safely drop a rung or two.

Two concrete triggers keep the rule honest. Escalate — a more capable model, or more effort — the moment outputs come back inconsistent between runs, thin on depth, or wrong on the hard cases in a small test set kept for exactly this purpose; de-escalate when a cheaper tier clears that same set without a wobble. And where the work is high-volume with a checkable failure signal, a third move beats both: run every task at low effort or on a cheaper model first, then re-run only the failures at the default. In Anthropic's measurements that pattern reached roughly the same quality as running everything at full strength for about half the total cost — the cheap pass does most of the work, and the expensive pass is spent only where it is actually needed.

A four-rung ladder: reach for the rung the work needs

Put roles on the model dial and it stops being a price ranking and becomes a set of jobs. Each rung earns its place on a different kind of work, and the per-token price on each is the sticker, not the total.

Exhibit — Four rungs, most to least capable: Fable opens the hardest problems, Haiku sweeps the checkable ones.. Reach for the rung the work needs; the per-token price on each is the sticker, not the total. Source: Roles per Anthropic's Claude Code model guidance (2026); relative token prices from Anthropic list pricing (Fable $10/$50, Opus $5/$25, Sonnet $2/$10, Haiku $1/$5 per 1M tokens), as of Sept 2026 — confirm current rates.

In my own week of building on these tools, the ladder maps cleanly onto the work. Fable comes out briefly, at the very start of a piece, for the genuinely wide-open thinking — mapping a sprawl of repositories and documents into one coherent thought, the one job where the top of the model dial reliably earns its keep because there is no clean spec to lean on. Once the shape of the problem is clear, the work hands off to Opus, kept near the top of the effort dial, for the building and the hard calls that fill most days. The cheaper rungs carry the repeatable, well-specified, easy-to-check work: a standing routine that sweeps open work each morning runs on Sonnet, and the reading-heavy sub-tasks that fan out — go read all of this and hand back the one line that matters — go to Haiku. That fan-out is where Haiku shines and also where its edges show: on factual reading verifiable at a glance it costs a fraction of the top tier, but its accuracy falls off a cliff on long agentic loops, so it is never handed one. High volume, clear instructions, cheap to check — that is its home, and it is a good one.

The prices on that ladder are the sticker, and they move — Anthropic's own tiers have been repriced more than once — which is exactly why they are dated here and worth confirming before anyone leans on them. The finished-task bill reorders that list constantly. Which points straight at the dial most people never touch.

The effort dial pays off on a different curve for different work

There is a common, hard-to-articulate experience among people who build with these tools: drop a capable model below its higher effort settings and it often stops doing what was expected, so the setting quietly gets turned back up. That is not superstition, and it is not the model having a bad day. It is that the effort dial pays off on a completely different curve depending on the kind of work — and building sits on the steep part of it.

The mechanism is intuitive once seen. Lower effort means fewer, more consolidated steps and less exploration. For a quick classification, less exploration is invisible; the answer is the same. For "map these repositories and reason about the architecture," less exploration is the degradation — the model stops looking around before it commits, and the result reads as a worse colleague. Anthropic frames the effort levels with illustrative curves that show exactly this split, and is careful to say the curves are for illustration, not measured benchmark data. The shapes are the point.

Exhibit — Turning up effort barely moves research, but it lifts building and deep work far further before it flattens.. Match the dial to the curve: crank it for building and deep synthesis; on research and chat, low effort costs almost nothing. Source: Conceptual shapes of Anthropic's effort-vs-quality curves ('Choosing a Claude model and effort level in Claude Code', 2026), which Anthropic states are illustrative — not exact figures.

Read the three shapes and the advice falls out of them. Research, knowledge work and chat sit on the flat line — turning effort down costs almost nothing there, so a lower setting is simply free money. Building, long-horizon and agentic work sit on the steep line, where a low setting gives up a meaningful chunk of success for a slice of the cost; this is why dropping below a high setting disappoints on building, and why the instinct to turn it back up is correct rather than a crutch. Deep, multi-part synthesis sits on the straightest line, where every step up buys real quality with no free cut anywhere. One honest complication rides on top: the gains flatten toward the ceiling, and effort is not free — a higher setting can generate on the order of seven times the tokens of a lower one for the same prompt. Some independent testing finds the very top setting adding cost for no accuracy, and occasionally regressing on specific coding tasks, which is why Anthropic recommends running the default effort for most work rather than pinning the dial to maximum. The practical read: keep effort high where the curve is steep, let it drop where the curve is flat, and treat "max everything" as the trap it usually is rather than a free upgrade.

A model is not one thing: performance is model × harness × product

The last correction settles more model debates than any benchmark. A model's real-world performance is not a property of the model alone. It is a property of the model plus the software it is wrapped in — what the field has started calling the "harness": the tools it can reach, how its context and memory are managed, its permissions, and how it recovers when it makes a mistake. Give the same model a better harness and its results move, and not by a little.

Exhibit — Same tasks, same models — swapping only the harness moves success by 23.8 points.. A leaderboard score was set in someone else's harness — treat it as a shortlist, then prove it in the one you use. Source: Harness-Bench (arXiv:2605.27922): 106 tasks, 8 model backends, 6 harnesses; best (76.2) vs worst (52.4) configurable harness under the same task set and model pool.

That is a 23.8-point swing in whether the job got done, driven by nothing but the scaffolding around the model — a gap as wide as the distance between adjacent models on a leaderboard. Claude Code is a harness: it reads files, runs commands, manages context and caching, holds permissions, and recovers from its own errors. The same model called raw through the API, with none of that, behaves differently; inside a different product again — a chat window, a design surface, an assistant that runs work unattended — it behaves differently a third time. Usage is even metered across those surfaces together, but the behaviour is not portable.

The harness even reaches back into how the model should be prompted. The newest models are tuned to handle ambiguity and to decide for themselves, so an over-prescriptive prompt written for an older, more literal generation can actively reduce the quality of the best one, crowding out the judgement it was built to exercise. The very instinct that makes a weaker model reliable — spell out every step — can hobble a stronger one. This is also why the harness is the part that cannot simply be rented: every competitor can license the same model at the same price and inherits each model improvement the day it ships, but the tools around it, the context and error handling, and the prompt conventions that fit this model in this product are built, not bought.

Two things follow. First, a benchmark score is a weak guide to what any given seat will actually get, because it was set in someone else's harness on someone else's tasks. As Ethan Mollick puts it, you would not hire a candidate on their exam results alone; the same "job interview" discipline applies to a model — build a few scenarios from real work, run them more than once, and judge the models head-to-head on the tasks that matter, blind to which produced which. Treat leaderboards as a shortlist, never a verdict. Second, do not carry a judgement across tools: "this model felt dull yesterday" may only mean it sat in a thinner harness yesterday. Every claim in this report is made about Claude Code specifically, and on purpose — because a model is not one thing, and the harness is half the result.

Bottom line

Almost every confused model choice dissolves the moment one number replaces another: stop optimising the price of the token, start optimising the cost of the finished task. The model dial sets the ceiling and the effort dial sets how much of it gets spent, and both are read off the task — its ambiguity, horizon, checkability, and volume — not off a price list. Start from a capable default, step down only where quality visibly holds, and step up the instant it slips; a cheaper model that flails is the expensive one, because task cost climbs with the square of the turns. Match the effort dial to its curve — high for building and deep synthesis, low for research and repetition, and rarely at the absolute ceiling, where returns flatten and tokens do not. And judge any of it where the work actually runs, because real performance is model × harness × product, and a leaderboard measured in another harness is a shortlist, never the answer. None of this is about spending more. It is about spending on the thing that was always the point — the finished job — and letting the price of the token go back to being what it always was: the sticker in the window, not the cost of the drive.

Sources