Skip to main content
Same prompt. Three engines. You are not looking for the “best model.” You are looking for the cheapest model that clears the bar for this job, and feeling what the others do differently. Plan for about 20 minutes. Work in claude.ai or Desktop. Use three separate new chats so earlier answers do not contaminate later ones.
Part 1

Read the brief and copy the prompt

3 min
Part 2

Run Haiku, then Sonnet, then Opus

12 min
Part 3

Score the three answers side by side

5 min

The job

You are the asset manager for Cedar Ridge Commons, a suburban retail center. Your CIO wants a tight recommendation before tomorrow’s IC: renew the anchor pharmacy early, or wait and re-trade later. This brief is designed so the models diverge. It has conflicting signals, a missing fact, and a hard format. Depth shows up as judgment, not as longer prose.

Part 1: Copy the prompt

Do not rewrite it between runs. The whole point is a controlled compare.

Model compare brief — Cedar Ridge pharmacy renewal

Part 2: Run the race

For each model, start a new chat, pick that model in the picker, paste the same prompt, and send. Do not enable research. Keep effort at the default unless your facilitator says otherwise. Keep all three answers open in separate tabs or chats so you can scroll them together in Part 3.
If your plan’s labels differ slightly, pick the fastest/cheapest model, the default mid model, and the deep-work model. Same experiment.

What “good” looks like

A strong answer, on any model, should roughly do this:
Hard format

Uses the five headings exactly, or close enough to scan in five seconds.

Missing facts

Flags that sales / occupancy-cost data was withheld, instead of assuming the store is healthy.

Real trade

Weighs the pharmacy ask against the clinic alternative and downtime/TI, not only against current rent.

IC-ready note

CIO note stays near 120 words and leads with the decision.

Part 3: Score the three answers

Tap every model that cleared each dimension. Ties are expected. Use this with your facilitator while the three chats are still open.
Then answer one sentence each:
  1. Cheapest model that cleared the bar: _______________
  2. What the deeper model bought you, if anything: _______________
  3. Default you will use for this kind of IC note next week: _______________

Haiku wins when the job is volume and the miss is cheap.

Sonnet wins for most analyst work you will still edit.

Opus wins when a shallow miss would survive into IC.

Haiku is often first and often soft on the withheld sales data or the clinic TI/downtime math. Sonnet usually clears the bar and is the practical winner for this job. Opus tends to sharpen the dual-track / refinance framing; if it only adds length, Sonnet still wins. If anyone’s Opus invents pharmacy sales comps, call that out as a failure, not depth.

Flight card

Debrief with your facilitator

  1. Where did the answers actually diverge: posture, risk, or just tone?
  2. Did Opus earn the upgrade on this brief, or did Sonnet already clear IC?
  3. Name one real task from your week that should drop to Haiku, and one that should start on Opus.