Skip to main content
Claude does not read words. It reads tokens: small chunks of text, usually three or four characters of English, sometimes a whole common word, sometimes a single punctuation mark. Every message you send gets sliced into tokens before Claude sees it, and every reply is assembled from them, one at a time. Tokens are also the currency. Usage limits, context windows, model pricing: all of it is denominated in tokens. Ten minutes of token literacy explains almost every limit you will ever hit.

Slice some text yourself

Type anything, or hit the presets. Watch where the knife falls.
What the slicer just showed you:
  • Common English words cost one token. The language Claude saw most during training is the language it reads most efficiently.
  • Rare words shatter. “Antidisestablishmentarianism” costs what a short sentence costs.
  • Numbers shatter too. A dollar figure with cents can cost more tokens than the sentence around it.
  • Other languages usually cost more tokens for the same meaning, and emoji are surprisingly expensive.
Rule of thumb: an English word is about 1.3 tokens, a page is about 400, a dense 60-page document is about 30,000.

Every turn prints a receipt

Here is the part almost nobody knows, and it explains almost every blown limit: when Claude answers, it re-reads the entire conversation first, and you are charged for all of it, every single turn. A six-message thread does not cost six messages. Run the printer and watch the “history, re-read” line.
The lesson is not “stop pasting documents.” It is: what enters a thread stays in the meter. Paste the two pages that matter instead of the sixty that exist, park reference material in project knowledge, and when a thread’s history stops earning its re-read cost, move to a fresh chat with a summary.

The two limits, side by side

People hit two different walls and call both “the limit.” They are different walls with different fixes.
Usage limitHow much Claude you can use across all chats and surfaces. Spent by every turn, everywhere, and refills on a timer.Hit it? Wait for the reset, work leaner, or upgrade.
Length limitHow big one conversation can get before the window is full. Never refills; a thread only grows.Hit it? Summarize and start fresh. The thread is done, you are not.
One connection between the walls: long threads attack both at once. The re-read receipt drains your usage limit while the growing history marches toward the length limit. Two footnotes worth knowing. Usage is shared across surfaces, so a heavy morning in Claude Code spends the same budget as your chats. And on paid plans with code execution enabled, Claude now manages very long conversations by summarizing older messages when the window gets tight. If you see it “organizing its thoughts” mid-thread, that is the window being cleaned for you.

Where usage actually goes

In rough order of damage:
  1. Immortal threads. The receipt above. The single biggest silent drain.
  2. Wholesale pastes. Sixty pages in a chat gets re-read forever; the same file in project knowledge is handled far more efficiently.
  3. Heavyweight settings on lightweight work. The most powerful model plus extended thinking, used to reformat a table. Match the tool to the job.
  4. Regenerate roulette. Rerolling a vague prompt five times costs five full turns. One sharper request is cheaper and better.
  5. Maximal output. “Write everything you know” costs tokens on the reply side too. Ask for the shape you actually want.