Watch the same question both ways
Two questions, two panes. The left pane answers cold. The right pane gets scratch paper. Run them and watch both the clock and the content. The discount question has a trap in the middle: two percentages that live on different clocks. The instant answer compared them anyway and sounded completely sure. The reasoning pane caught the trap on its third line. The subject-line question has no trap, so the scratch paper bought a few seconds of nothing. That is the entire decision rule. Thinking pays on problems with a trap in the middle, and costs time and tokens everywhere else.Worth it, not worth it
Flip it onAnything with numbers that interact: rates, dates, discounts, capacityMulti-step logic where step 3 depends on getting step 1 rightPlans and trade-offs with constraints that can conflictReviewing work for errors, yours or Claude’sAnything you would personally do on scratch paper
Leave it offLookups and simple factual questionsRewrites, tone changes, formattingBrainstorms where volume beats depthQuick back-and-forth where latency is the experienceAnything you will fully verify yourself anyway
What thinking is not
Three misconceptions do most of the damage:- It is not a knowledge upgrade. The model knows exactly as much either way. If the answer needs fresh facts or your files, thinking harder about missing information produces confident fiction with extra steps.
- It is not a truth guarantee. Reasoning can be fluent and wrong. The visible trace makes errors findable, not impossible.
- It is not free. Reasoning tokens cost time and money like any others. Thinking on everything is not diligence, it is paying for scratch paper at the sandwich counter.
The model already has opinions
Current Claude models decide on their own whether a request deserves deep reasoning, a judgment call the docs name adaptive thinking. Ask something simple with thinking on and you will often see a one-line trace that says, in effect, no trap here, answering directly. Ask something genuinely hard and the model digs in without being told. Two settings interact with this and are worth keeping straight. Effort controls how thorough the whole response is, and it is a separate dial from the thinking switch: any combination is valid. And on Opus 5 the switch will not turn off at all, because that model always reasons before answering.Read the trace like a reviewer
The expandable reasoning section is not decoration. It is the closest thing you get to watching a colleague’s process, and it makes one skill possible that no other feature offers: diagnosing where a wrong answer went wrong. When a thinking answer misses, open the trace and find the first false line. Nearly always it is one of two things: the model misread your intent, which is a context problem you fix by rewriting the request, or it made a bad assumption on a detail you never specified, which you fix by specifying it. Either way, the retry is surgical instead of “try again and hope”.What tool use is
Thinking helps Claude reason about what it already has. When the answer needs facts from outside its head, that is a job for tools.