Skip to main content
Back to Course 102
3/13
Unit 3 of 13
Unit 02 — Phase 01

The Right Brain for the Job

Configure the Machine

Listen to this unit

Read-aloud narrates the unit in a natural voice, paragraph by paragraph. It needs an account.

In short

Thinking is a dial you rent. Model tiers and deliberation modes are staffing decisions, and the rates are on a public menu. Route with two questions: how many dependent steps, and how bad is wrong. Default cheap, escalate on measured gaps, rent the frontier only where both answers are high.

Socratic Mode

The Hook

Every few weeks, a new model launches with a chart showing it beating everything else. Your feed fills with people declaring the old tools dead. And here's the strange part: when you actually try the new model on your own tasks, half the time you can't tell the difference.

That's not because you're missing something. It's because at the top, the models really are converging on everyday tasks, while their prices differ by ten times or more, and their behavior on hard, multi-step problems differs in ways no launch video shows you.

Most people resolve this by picking one tool and paying whatever it costs. Operators do something else: they treat models like a staffing decision. Different brains, different rates, different jobs. This unit gives you the selection method, and a way to test any new model in thirty minutes flat.

Video — The Dial You RentWatch on YouTube

The Core Concept

Here's the mental model: thinking is a dial you rent.

Modern assistants don't offer one kind of intelligence. They offer a range, and you pay along two axes at once. The first axis is the model tier: most providers ship a fast cheap model, a balanced middle model, and an expensive model. The second axis is the : many models can either answer immediately or deliberate first, working through a problem step by step before responding. More deliberation costs more time and more money. On some problems it buys you nothing. On others it's the entire difference between a wrong answer and a right one.

Think of it like hiring. You don't send a senior specialist billing by the hour to sort the mail, and you don't hand the merger contract to the intern because the intern is cheap. The skill isn't loyalty to one person on the team. It's matching the brain to the job, and knowing the rates. With AI, the rates are printed on a public menu and the specialists are available in seconds, which makes mismatching them the only remaining way to get this wrong.

Explained your way: the AI rewrites this idea around something you already know.

To match well, judge any model on five dimensions: capability (how hard a problem it can actually solve), deliberation (whether the thinking dial is available and how far it turns), speed (seconds versus minutes matter when a task repeats daily), price (per use, which you'll learn to estimate), and fit (context size, file and image handling, and whatever your task specifically needs). Launch marketing talks almost entirely about the first dimension. Your costs live in the other four.

Then run the four-step lineup for any task:

  1. Default to the cheap tier. Most daily volume is extraction, reformatting, summaries, and simple drafting. The cheap tiers handle these indistinguishably from the frontier.
  2. Escalate on measured gaps, not vibes. If the cheap tier's output fails your check, move up one tier and rerun. A gap you observed is a reason; a launch video is not.
  3. Turn the thinking dial for step-heavy problems. Scheduling with constraints, multi-stage math, planning with dependencies, anything where step 4 depends on step 2. Deliberation is the product there.
  4. Rent the frontier only where the gap justifies the rate. High stakes plus hard problems. Know what it costs per run, because the next units multiply every run.

The routing decision compresses into two questions you can ask about any task. How many dependent steps does it have? And if it's wrong, how bad is that, and how fast can you check? Run your own real tasks through it:

Route Your Own Tasks
Unit 02 · The Right Brain for the Job

Enter up to five real tasks from your week. Answer two questions per task and the grid routes each one to a tier, a verification level, and a cost band.

How many dependent steps?
If it's wrong, how bad, and how easy to check?

Low steps and low stakes route cheap. High steps buy deliberation. High stakes buy verification, which is a habit, not a model tier. Only the corner where both are high rents the frontier.

Stanford's AI Index measured one of the fastest cost collapses in the history of technology: querying a model at the level of GPT-3.5 (the original ChatGPT brain, state of the art in late 2022) fell from about $20 per million to about $0.07 in roughly two years, a drop of about 280 times. The pattern has continued: intelligence that costs frontier prices today becomes nearly free within a couple of years, because every provider's mid tier keeps absorbing last year's frontier.

Two operator conclusions follow. First, yesterday's "premium" capability is almost always available today at commodity prices, so defaulting to the cheap tier is not settling; it's arbitrage. Second, any specific model or price this page could print will be wrong soon, which is why this unit teaches you a measurement method instead of a shopping list, and why the live menu below matters more than any article.

In April 2025, Meta released Llama 4, and one of its models briefly ranked near the top of LMArena, a popular leaderboard where humans vote between model answers. It then emerged that the version submitted to the leaderboard was a special variant tuned to please voters, different from the model actually released, and the leaderboard's operators changed their policies in response.

The lesson isn't that one company misbehaved. It's structural: any public benchmark becomes a target, and targets get gamed, legitimately and otherwise. Leaderboards are useful weather reports, and independent trackers that measure price and speed alongside quality are better ones. But the only benchmark that can't be gamed against you is the one made of your own tasks, run by you.

Which brings us to the instrument this unit exists to give you: the personal benchmark. Pick five or six tasks you actually repeat: one extraction, one summary of your kind of document, one draft in your voice, one step-heavy problem from your life, one task from your Unit 01 domains. Save the exact inputs. Write one line per task defining what "good" looks like. Store the whole thing in a Unit 01 workspace. Total build time: about thirty minutes, once. From then on, every noisy launch week becomes a quiet half-hour test: run the six inputs through the new model, compare against your current results, decide with numbers.

Knowledge check
Sara has three tasks today: extract emails from a messy document, plan a study schedule with eight overlapping constraints, and draft a two-line reply to a friend. Using the two-question grid, the best routing is:
Knowledge checks save to your account.

Live Demo

Free path: every step below works on free tiers. Free plans usually let you switch between at least two model options and a thinking toggle; if a step's exact control is paid-only in your app, run the comparison across two different free assistants instead.

Step 1, the task that doesn't care. Paste any messy text containing a few email addresses into your assistant's cheapest and most expensive available settings:

Prompt
Extract every email address and the organization it belongs to. Two-column table.

Identical results, most likely. Remember the feeling: most of your daily volume is this step.

Step 2, the task that cares a lot. Give the model a genuinely constrained problem with thinking off, or on the fastest setting:

Prompt
Build me a weekly schedule: football practice Mon/Wed 5-7pm, part-time shifts Tue/Thu 4-8pm, at least 6 hours of study spread over at least 3 days, one full rest evening, gym 3x for an hour, never gym and football the same day. Produce the schedule, then verify every constraint one by one.

Now run it again with thinking on. Compare the schedules, and especially the verification step. This is the class of problem where deliberation is the product.

Step 3, the Studio's routing. The Studio runs both kinds of task: captions (short, checkable, daily) and the weekly content strategy (many dependencies, higher stakes). Route them with the grid: captions go cheap; strategy gets the dial. That one routing decision will cut the Studio's running costs several times over once it's automated in Unit 08.

Step 4, check the independent menu. Look up the models you just used and compare measured quality, speed, and price against what the launch marketing claimed:

Interactive tool — Artificial Analysis: live independent model comparisons

Artificial Analysis: live independent model comparisons

This tool opens in a new tab where you can interact with it directly.

Launch Tool

Step 5, benchmark v1. Draft your personal benchmark from the recipe above: five or six tasks, exact inputs saved, one line each on what "good" looks like, stored in a Unit 01 workspace. You'll use it in the very next unit, and in every unit after that.

AI Cost Calculator
Unit 02 · The Right Brain for the Job
Pricing (Claude Sonnet)
Input: $3.00 / 1M tokensOutput: $15.00 / 1M tokens

Adjust the sliders to see how compute costs scale. Every message you send costs real money.

Queries per user per day50
Input tokens per query (your prompt)500
Output tokens per query (AI response)800
Users1,000
$675.00
Per day
$20.3K
Per month
$243.0K
Per year
$20.25
Per user / month
The model loyalist
One tool for everything, chosen by habit or hype. Pays flagship prices for extraction and summaries, and never learns that the mid tier overtook their flagship months ago, because nothing in their workflow would surface it.
The portfolio operator
Cheap tier by default, thinking modes for step-heavy problems, frontier only where a measured gap justifies it. Runs a 30-minute personal benchmark on any model worth considering, and lets results decide instead of launch videos.

Operator Moves

Start cheap, escalate on gaps. Make the cheap tier your reflex for everything. Escalation requires evidence: a failed check, not a feeling. You'll be right far more often than the person doing the reverse, and you'll spend a fraction as much.

The launch-week ritual. When a new model drops, don't read the takes. Run your six benchmark inputs through it, compare against saved results, and decide in thirty minutes. You'll evaluate faster than the commentators and more accurately, because your evidence is yours.

Price the run, not the month. Before adopting any model for a repeating task, estimate cost per run and multiply by frequency. A tier that's 5x too expensive on a daily task costs you that, times thirty, forever. This number is why "start cheap" is a rule and not a vibe.

Why This Matters

Right now the stakes are a few dollars and a few seconds. But this course is headed somewhere specific: research agents that make dozens of model calls per question, coding agents that run for an hour, automations that fire daily while you sleep. In those settings, model choice multiplies. A tier that's five times too expensive or twice too slow costs you that, times every step, times every run, times every day. The selection habit you build at chat scale is what keeps agent-scale bills sane later.

The subtler payoff: the specific models in this unit will be museum pieces embarrassingly soon, which is exactly why the "Last updated" stamp sits at the foot of this page. What doesn't depreciate is the practice. Tiers matched to task structure, escalation on measured gaps, and a benchmark that turns every noisy launch week into a quiet half-hour test. Tools expire. Instruments don't.

Knowledge check
Stanford's AI Index found the cost of GPT-3.5-level performance fell from about $20 to about $0.07 per million tokens in roughly two years. The most useful operator conclusion is:
Knowledge checks save to your account.

The Challenge

Benchmark and Policy

40 minutesHands-on

Build the instrument you'll use for the rest of this course, and the policy it feeds.

  1. Build benchmark v1: five or six real, repeating tasks. Save the exact input for each and one line defining "good." Store it in a Unit 01 workspace.
  2. Run it twice: once on a cheap tier, once on the most capable setting you can access for free (a higher tier, or thinking mode on). Score each task pass or fail against your "good" line.
  3. Find your gaps: for every task where the cheap tier failed and the capable setting passed, you've found a measured gap. For every task where they tied, you've found free money.
  4. Write your model policy (5 lines or fewer): which tier is your default, which task types get the thinking dial, what evidence triggers escalation, and what a frontier task looks like for you.
  5. Price one repeating task: using current prices from Artificial Analysis or your provider's pricing page, estimate the monthly cost of your most frequent task on two different tiers.
Success criteria: a stored, rerunnable benchmark; at least one measured gap or measured tie you can name; and a written policy specific enough that a friend could route your tasks with it.
Submitting your work needs an account.

Key Takeaways

  1. 1Thinking is a dial you rent. Model tiers and deliberation modes are staffing decisions, and the rates are on a public menu.
  2. 2Route with two questions: how many dependent steps, and how bad is wrong. Default cheap, escalate on measured gaps, rent the frontier only where both answers are high.
  3. 3Prices collapse and leaderboards get gamed, so the durable instrument is a personal benchmark of your own tasks: thirty minutes to build, thirty minutes to rerun on any launch day.
  4. 4Model choice multiplies once agents and automations enter. The selection habit you build now is what keeps those bills sane later.

The Rabbit Hole

Type: Report Title: Stanford HAI AI Index, the annual state of AI URL: https://hai.stanford.edu/ai-index Description: The measurement culture this unit teaches, at planetary scale: capability, cost curves (including the price collapse in this unit), and adoption, updated every year. Skim the ten headline charts; they're the industry's benchmark v1.

Explore Further

TypeTitleURLDescription
ReportStanford HAI AI Indexhai.stanford.edu/ai-indexAnnual measurements including the ~280x inference price collapse
ToolArtificial Analysisartificialanalysis.aiLive independent comparisons of model quality, speed, and price
ArticleThe Verge, "Meta gamed the system with Llama 4 benchmarks" coverage (Apr 2025)theverge.com/meta/645012/meta-ll…The LMArena episode behind this unit's leaderboard warning
ToolAnthropic pricingclaude.com/pricingCurrent Claude plan and usage pricing
ToolOpenAI pricingopenai.com/pricingCurrent OpenAI plan and usage pricing
ToolGoogle AI plansone.google.com/about/google-ai-pla…Current Gemini tiers and what each unlocks
DocsAnthropic, extended thinking documentationdocs.claude.com/en/docs/build-with-…How a deliberation dial works under the hood, from one provider

Last updated: August 9, 2026. Model names and prices in this unit expire fast, which is exactly what this unit is about. The selection method doesn't.

Track your progress

Marking a unit complete, the Prove It check and your place in the course all need an account. The reading stays free.