Skip to main content
A language model on its own can do exactly one thing: write. It cannot check today’s news, cannot multiply large numbers reliably, cannot open your spreadsheet, cannot save a file. Everything it produces alone is prediction, polished to sound like knowledge. Tool use is the escape hatch. Mid-answer, Claude can stop writing, call something real, read what comes back, and continue with facts it did not have a moment ago. A search, a fetched page, a program it wrote and ran, a file it built. Each call leaves a visible record in the conversation. That record is the point of this lesson. An answer built on tools comes with receipts, and you can audit every line.

One question, with and without receipts

Same request on both sides: what did our competitor announce this week, and how should it change our pricing pitch? The left side answers from memory. The right side is allowed to work. Print the receipt and inspect any line of it.
Notice what the receipt exposed. The memory answer invented a price cut, because a price cut is the kind of thing competitors announce and prediction fills gaps with the plausible. The real announcement was a distribution deal. Nothing about the memory answer sounded wrong. It was just never connected to the world.

The four built-ins

Claude ships with a toolbox that is already switched on. No setup, no configuration, available in the app the day you get an account:
Web fetchOpens a specific page you point at and reads the whole thing, not a summary of a snippet of it.
Code executionA private computer inside the chat. Claude writes real code and runs it: exact math, data analysis, file crunching. Arithmetic stops being prediction.
File creationBuilds actual files you can download or send to Drive: spreadsheets, decks, documents, PDFs. Output that opens in the tools your colleagues use.
Beyond the built-ins sit connectors, which extend the same mechanism into your own systems: Drive, Slack, your CRM, your wiki. Different plug, same principle. Every one of them turns a category of guessing into a category of checking.

Say the tool’s name

Claude decides on its own when to reach for a tool, and it usually decides well. But on work where correctness matters, do not leave it implicit. Ask for the act, not just the answer:
  • “Search for this week’s coverage before answering” instead of “what happened this week”
  • “Compute this with code and show the code” instead of “what do these numbers work out to”
  • “Read the page at this link” instead of pasting a link and hoping
  • “Build this as a spreadsheet” instead of accepting a table in chat you will retype anyway
The phrasing sounds fussy exactly once. After that it is just how you talk to a system that can either predict or verify, and you are choosing which.

The auditing habit

One behavior separates fluent users from everyone else: before trusting an answer, they glance at what ran. Tool calls are visible in the conversation as they happen. If an answer contains a number, a quote, a date, or a claim about the current world, and no tool ran, treat that piece as a guess wearing a suit. It might be right. It earned nothing.
Turn this into a reflex with one question per answer: where did that come from? If the answer traces to a search result, a fetched page, or code output, trust accordingly. If it traces to nothing, verify or ask Claude to verify, which usually just means asking it to use the tool it skipped.

The agentic loop

One tool call answers a question. Chain many of them behind a goal, with the model checking its own results, and you get agentic work.