Skip to main content
Claude is not one brain running at one speed. Behind the chat box sits a family of models, a dial for how hard the chosen one works, and a switch for whether it reasons out loud before answering. All three controls live in one small menu next to the send button. Most people never open it, which means most people run every task, from “fix this typo” to “rebuild the pricing model”, on the exact same settings. That is like owning a car with one gear.

Sit at the desk

Here is that menu, blown up into a control desk you can actually play with. Pick an engine, slide the effort, flip the thinking switch, and watch the three meters respond.
Two quirks you just discovered by touching the desk are real product behavior. Haiku has no effort fader, it runs at one intensity. And on Opus 5 the thinking switch will not turn off, because that model always reasons before it answers.

Meet the four engines

Every model in the family is a different answer to the same trade: depth against speed against cost. None of them is “the best”. Each is the best at a different shape of work.
Fable 5The frontierThe most capable model Anthropic ships. Reserved for problems where Opus demonstrably falls short, because you pay roughly double for the margin.
Opus 5The deep workerNear-frontier intelligence at half the frontier price. Long documents, hard analysis, work where a wrong answer costs real money. Always thinks before it speaks.
Sonnet 5The daily driverThe default for a reason. Plans, uses tools, and handles the everyday mix of drafting, analysis, and coding at a fraction of Opus cost.
Haiku 4.5The sprinterFastest and cheapest by far. Built for volume: hundreds of small, similar items where latency and cost matter more than depth.
Hold the cost ratio loosely and it explains almost every model decision: if Haiku costs one unit, Sonnet costs about two, Opus about five, and Fable about ten. Depth is priced honestly. Spend it where depth is what the job needs.

Modes are the other half

Which model you pick is only the first control. The same engine behaves very differently depending on the mode it runs in:
  • Effort sets how thorough every response is, from a quick pass to an exhaustive one.
  • Thinking gives the model drafting space to reason before it commits to an answer, and lets you read that reasoning.
  • Tools let the answer touch reality: search the web, run code, build files, reach your systems.
  • The agentic loop chains all of it, letting Claude plan, act, check its own work, and go again until the goal is met.
Each of these gets its own lesson in this section, with the same rule throughout: settings are per-message, not per-lifetime. You can change any of them mid-conversation, and the next reply obeys.

What’s in this section

Choosing the right model

One brief, four engines, four different bills. How to pick without a chart.

What thinking is and is not

The toggle that buys drafting space, what it catches, and what it costs.

What tool use is

How an answer earns its receipts: search, fetch, code, and files.

The agentic loop

Plan, act, observe, revise. Why agents lap the problem instead of guessing once.

Skills, MCP, connectors, and built-ins

Four words people mix up daily, shown as a simple stack.

Lab: Same brief, three models

Run one IC renewal brief on Haiku, Sonnet, and Opus. Score format, judgment, and which engine actually clears the bar.

Quiz: Models & modes

Ten instructor-led questions before you leave this section.