Skip to main content
The model picker is the highest-leverage click in Claude. Same prompt, same context, different engine: the answer changes in depth, arrives at a different speed, and costs a different amount. People who never touch it are paying deep-work prices for typo fixes, or asking a sprinter to write board strategy. The way to learn the trade-off is not a spec sheet. It is watching all four models run the same job.

Race them

Pick a brief, start the run, and watch the lanes. First across the line is not the winner. The winner is the cheapest engine that clears the quality bar for that job, and the bar moves with the job.
Run all three briefs and the pattern falls out on its own. The volume job went to the cheapest engine because the bar per item was low. The board memo went to Opus because the bar was high and Fable’s extra margin bought nothing Opus had not already delivered. The daily mix went to Sonnet because most work lives in the middle.

Three questions replace the chart

You do not need to memorize benchmarks. Before a task, ask:
1

What does a miss cost?

If a mediocre answer costs you thirty seconds and a retry, go cheap and fast. If it goes to the board, misprices a deal, or ships to a customer, buy depth. Price the failure, not the prompt.
2

How many of these are there?

One memo is an Opus job. Four hundred ticket summaries is a Haiku job even though each summary is easy, because cost and latency multiply by volume while depth does not.
3

Will a human check it?

Work you will read closely anyway can afford a cheaper first draft. Work that goes out on trust needs the model you would trust unsupervised.

The escalation habit

Fluent users do not deliberate per task. They run one default and move off it on evidence:
  • Start on Sonnet 5. It clears the bar for most knowledge work, and you will know within one reply when it has not.
  • Escalate to Opus 5 when the reply misses structure, drops constraints, or the stakes make a miss expensive. Opus was built to deliver near-frontier quality at half the frontier price, so it is the natural ceiling for daily escalation.
  • Reach for Fable 5 only when Opus has demonstrably fallen short on this specific problem. Paying double to confirm Opus was already right is the most common way to waste the frontier.
  • Drop to Haiku 4.5 the moment a task becomes a conveyor belt: many small items, same operation, speed felt on every one.
The picker works mid-conversation, and the change applies from the next reply. A common power move: set up the problem and gather context on Sonnet, switch to Opus for the one hard synthesis step, then switch back for the follow-ups. You pay deep-work prices for exactly one message.

Where the defaults come from

Claude picks a sensible default for your plan, and the menu next to the send button (“More models” expands the full list) is where you override it. The default is a starting posture, not a recommendation for the task in front of you. The habit above is the recommendation.

What thinking is and is not

Model choice sets the engine. The thinking toggle decides whether that engine gets drafting space before it answers.