Price and Pace

Two September 2026 publications try to put numbers on how fast AI is moving. Epoch AI measured how fast the price of a fixed level of AI performance is falling. Anthropic proposed how to measure, from inside a lab, how much of AI development AI is now doing itself. They look like different subjects. They are the same kind of instrument, with the same kinds of weakness, and one quietly depends on the other. This course teaches both, then puts them together.

7 modules
2 interactive tools
14 reasoning questions
~60 min

Built on The plunging price of thought (Epoch AI, 22 Sep 2026) and Measurements for understanding the pace of AI development inside frontier labs (Anthropic, Sep 2026)

How to read this course

Both publications are careful about their own limits. Many of the sharpest caveats below are the authors’ own, and where that is so the course says so. Objections that go beyond what the authors concede are in Module 7 and are labelled as objections.

Short on time? Do Module 1 (the shared structure) and Module 6 (where the halves meet).

Course Modules

  1. Two measurements, one problemStart here
  2. How to price a thoughtEpoch: method
  3. What the price result says, and does notEpoch: results
  4. How much of AI R&D is AI doing?Anthropic: automation
  5. Watching the agents, counting the computeAnthropic: oversight
  6. Where the halves meetInteractive
  7. Objections, and what each number licensesInteractive
1

Two measurements, one problem

Both are index numbers, and index numbers fail in predictable ways
By the end of this module you will
  • See why “price of thought” and “share of R&D done by AI” are the same kind of measurement
  • Know the four design choices every such measurement makes, and where each can go wrong

You cannot measure “thought” or “AI development” directly. Neither exists as a unit. So both publications do what statisticians do with inflation. They define a basket of things that can be measured, decide on a level to hold fixed or rate, choose weights to combine them, and choose a judge to score each item. The result is an index. Epoch says so explicitly: by averaging across benchmarks and performance levels, “we are effectively chaining price drops for various services. We are not tracking a conceptually unitary ‘price of thought.’”

Design choiceEpoch: price of thoughtAnthropic: R&D Automation Index
BasketFive primary benchmarks: FrontierMath T1–3, OTIS Mock AIME, GPQA Diamond, Chess Puzzles, a secret Mystery GameA frozen tree of 378 kinds of R&D task, built from July 2026 work records
Level held fixed / ratedA fixed accuracy score; measure the cheapest cost to reach itAn automation level, AL0–AL5, for each task
WeightsAveraged evenly over a grid of dates and accuracy levelsPerson-time: each sampled person’s week split evenly across their tasks
JudgeThe benchmark’s own answer keyAn independent Claude judge, checked against staff ratings
Blind spotBenchmarks are not useful work; nobody actually buys at the frontierA frozen basket cannot see new kinds of work; Claude is judging Claude
The reason to study them together

Index numbers have a small set of known failure modes: the basket drifts away from what matters, the weights encode a choice, the judge shares the errors of the thing judged, and a flat share can hide a changing quantity. Every one of them appears in these two publications. Learn to spot them once and you can read any AI-progress metric you will be shown next year.

Takeaways
  • Both are index numbers: basket, level, weights, judge
  • Each design choice is also a claim about what matters
  • The same four failure modes recur across both
2

How to price a thought

Why price per token stopped working, and what Epoch measured instead
By the end of this module you will
  • Explain why reasoning models broke price-per-token comparisons
  • Describe the cost frontier and the truncation trick that maps it

Why price per token stopped working

Earlier estimates compared the per-token price of models that crossed a score threshold. Reasoning models broke that: they “can productively consume far more tokens, but as a result extract good performance from an underlying LLM that is smaller and cheaper to run per token.” A cheap-per-token model that thinks for ten times as long is not cheap. So Epoch measures the total cost to reach a score.

The cost frontier

For each benchmark, every run gives a point: accuracy, release date, dollar cost. The cost frontier C(t, a) is the cheapest cost at which any model released by date t reaches accuracy a. The question is how fast that frontier falls while a is held fixed. It is the price of a fixed quality, the way economists track the price of computing power rather than the price of “a computer”.

The truncation trick

Mapping a model’s whole cost-versus-accuracy curve would mean running it at many budgets. Instead Epoch follows a method from the federal Center for AI Standards and Innovation (CAISI). Take one long run’s transcript. For any per-question budget X, count only the questions answered correctly within X output tokens, and score the rest at the chance-guess rate (0.25 on four-way multiple choice). One run then yields a whole curve.

The obvious worry is that a model told its budget might do better than one silently cut off. Epoch tested this. Announcing a budget helped high-effort models at low budgets, but they did not beat low-effort variants run normally. So running the method across all thinking levels covers most of the achievable territory. Epoch also filled a gap it found in its own data: past benchmarking had chased peak scores, so it ran extra low-effort and small-model evaluations to avoid biasing the frontier.

No standard errors, on purpose

A frontier is an extreme: one missing or extra run can move it for months. Sampling was “a bit ad hoc”, because Epoch mostly tested models near the frontier. And the grid is autocorrelated. So the authors write that they “abandon the formal paradigm of statistical inference”: they report reasonable but rough measurements, not confidence intervals. That candour is worth noticing. The right way to read their numbers is as a size, not a decimal.

Takeaways
  • Measure cost to reach a score, not price per token
  • Track the frontier: the cheapest way to reach each score at each date
  • The numbers are deliberately rough; read them as magnitudes
3

What the price result says, and does not

47% a quarter, why it varies, and the four caveats the authors attach
By the end of this module you will
  • Convert between quarterly and annual decline rates without error
  • Explain why the decline is fastest right after a score becomes state of the art
  • Name the limitations and which one matters most for real users

The headline

“The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year.”Luke Emberson and David Roodman, Epoch AI

The worked example: OpenAI’s o3 (January 2025) reached 75% on GPQA Diamond for about 30 cents a question. GPT-5.6 Luna, just under 18 months later, matched it for $0.0004. The authors call it “a 725-fold drop in the price of thought in under 18 months”, like a $50,000 car falling to $69. By comparison, US residential electricity fell about 1.05× a year (1892–1973), lithium-ion batteries 1.16×, compute 1.51× and DNA sequencing 1.84×.

47%/qtraverage, five primary benchmarks (≈13×/yr)
50–52%math, per quarter
39–43%game puzzles, per quarter
27.5%SWE-bench Verified (coding), per quarter
66%per quarter at the moment a score is state of the art
32%per quarter two years later

Why it falls fastest at the frontier

For three of the five main benchmarks, cost falls fastest right after a score first becomes state of the art. The authors’ explanation: the lab that gets there first can briefly charge a premium, then competitors, open and closed, catch up and the price plunges. Their implication is pointed: “supra-normal profits in LLM service provision may be fleeting for any given model,” which they call “surely both a consequence and a cause of the race dynamic.” Hold on to that line. It links price to pace in Module 6.

The authors’ own caveats

  • Benchmaxxing. Labs may train for known benchmarks. The secret Mystery Game falls a bit slower (44.0% model-free against 47.0%), which they read as “consistent with the benchmaxxing critique being valid but not fatal.”
  • Benchmarks are not useful work. Cost to score is a proxy for the market value of inference, not a measure of it.
  • Nobody lives on the frontier. “Essentially no user stays permanently on the cost frontier.” Real users switch models rarely and capture less of the decline.
  • Averaging choices matter. Other reasonable ways of building the grid give 42.9–58.0% instead of 47.0%.
Takeaways
  • ~47% a quarter compounds to ~13× a year; coding is closer to 3.6×
  • Fastest near the frontier, which means any model’s premium is short-lived
  • Real users capture less than the frontier decline
4

How much of AI R&D is AI doing?

The automation scale, the index built on it, and the boundary that carries the headline
By the end of this module you will
  • Know the automation levels and the difference between “collaborates” and “leads”
  • Explain how the index was built and validated
  • Locate the exact boundary the 26% depends on

Why Anthropic wants this number

The post opens: “AI systems are getting more powerful, and they’re increasingly being used to build the next version of themselves.” Its stated aim is to show how close the world is to recursive self-improvement, “a model fully autonomously building its successor”. It is framed explicitly against Dario Amodei’s call to pace the frontier: Anthropic says it “would expect these numbers to shift if there were coordination on pacing.”

The scale (from Epoch AI)

AL0
No AI involvement
AL1
Minimal AI involvement
AL2
AI assists
AL3
AI collaborates: it can do large chunks of work under close human direction
AL4
AI leads: it can complete most of the task end-to-end from a high-level prompt, while the human supervises
AL5
Fully autonomous, no human in the loop. “A level we have not yet reached.”

Anthropic’s own example makes the AL3/AL4 line concrete. A nightly data pipeline breaks. At AL3 the engineer brings the logs, stays tuned in, and Claude stops when something unexpected turns up. At AL4 the engineer hands over the alert. Claude finds the cause, fixes and tests it, handles surprises itself, reruns on a copy of the data and writes it up, then tags the engineer, who decides whether it ships. At AL5 nobody would have had to notice the failure at all.

How the index is built

  1. Catalogue the work. For each week of July 2026, sample 20% of staff in each department in the model R&D loop. A Claude research agent reviews each person’s week from Slack and internal documents. That yields about 15,000 granular tasks, organised into a tree of 542 nodes, 378 of them leaves (for example “serving incident postmortems”). The tree is then frozen.
  2. Rate each node. A Claude agent researches how that work is done; an independent Claude judge assigns a level, seeing only evidence from that month or earlier.
  3. Weight by person-time (Module 1).
  4. Validate. Staff who own each area rated it blind. Model and human agreed exactly 59% of the time, against 35% for two humans, and were within one level 97% of the time.

What it found (August 2026)

26%of AI R&D work where Claude “leads” (AL4)
>90%at or above “collaborates” (AL3)
0%fully autonomous in any measured subset

The accompanying chart shows the AL4 share rising from under 1% in February 2026.

Where the headline lives

The 26% counts tasks rated AL4 rather than AL3. That boundary is exactly where Anthropic says ratings are softest: “There remains real room for disagreement on borderline cases, such as where exactly ‘AI collaborates’ ends and ‘AI leads’ begins.” Within one level 97% of the time is reassuring about the index overall. It says less about a number defined by a single level boundary. The 26% is a real signal, and its precision is loose.

Anthropic names the two obstacles to comparing labs: there is no common methodology, and the “judge” model “could make the same kinds of errors as the model it is checking.” Its proposed fixes are third-party verification or other developers’ models as judges. It also suggests these measures could become triggers, “like a fixed testing window before a new model is used for further AI R&D.”

Takeaways
  • AL3 collaborates; AL4 leads; AL5 not reached
  • 26% at AL4 and >90% at AL3 or above, as of August 2026
  • The headline sits on the softest boundary in the scale, and the basket is frozen
5

Watching the agents, counting the compute

Coverage, latency and escalation for 30,000 agents, and a one-week compute snapshot
By the end of this module you will
  • Name the three oversight metrics and compute what the reported rates imply
  • Explain why a compute share can move without safety effort moving

Oversight

About 30,000 agents did research and engineering work at any one time on Anthropic’s most-used internal platform in August 2026. The reason to measure: “individual agents rarely misbehave. But when there are millions or billions of agents operating in the economy, even rare events can happen regularly.”

Online monitorOffline monitor
ForFast, irreversible harms (for example copying weights out)Slow, reversible harms (for example signs of misalignment)
Coverage100% of actions checked before execution, usually within seconds100% of actions ingested after the fact
LatencyAutomated review before action; humans review blocked actions within a week~100,000 flagged transcripts a week triaged by classifiers; ~50 a week escalated to humans
Rate0.002% of over a billion August decisions blocked (about 1 in 47,000)1–2 transcripts in every thousand flagged

Two design choices sit behind the numbers. Each agent has a persistent identity that survives model upgrades, so every action is auditable per agent. That lets an agent “treat what comes from another agent as a claim to check rather than a thought of its own.” And agents talk on an open, shared message system, cross-referenced with transcripts, so monitoring can follow communication between agents.

Compute

For one week in July 2026, Anthropic classified its compute. About 6% of compute going to AI R&D went to safety, and about 12% of compute going to AI-driven AI R&D. The estimates are “deliberately conservative”: work that helps capability as much as safety counts as R&D, and safeguards classifiers, “a separate, comparable amount of compute”, are excluded. The method: a Claude classifier reads each run’s metadata and code, on a 14% sample of the week’s ~10,000 runs, weighted toward the biggest. It agrees with human reviewers within a point or two.

The authors’ most important caveat is theirs, not a critic’s: “compute share measures only what is spent. A more efficient safety classifier, or a faster inference stack for production models, lowers the safety portion, but doesn’t mean we’re doing less safety work.” Module 6 puts a size on that sentence.

Takeaways
  • Oversight = coverage, latency, escalation; 100% automated coverage, thin human layer by design
  • Tiny rates on huge denominators are large counts
  • Compute share measures spend, not effort, and the authors say so themselves
6

Where the halves meet

Cheap thought drives the pace, and changing prices distort the pace metrics
By the end of this module you will
  • Explain two concrete ways the price of thought feeds into the pace of AI development
  • Show, with Epoch’s own spread of rates, how a flat compute share can hide a large shift in work

Link 1: cheap thought is what makes an agent workforce possible

Thirty thousand agents running continuously is an inference bill. When the cost of a fixed level of performance falls about 13× a year, a workforce of agents at a given capability gets 13× cheaper to run each year, or 13× larger for the same money. Epoch measures the price; Anthropic measures the result: the share of R&D where AI leads went from under 1% to 26% in six months. Neither publication cites the other, but read together they describe one mechanism.

Link 2: fleeting premiums are a race engine

Epoch finds prices fall fastest just after a performance level becomes state of the art, so “supra-normal profits … may be fleeting for any given model”, which the authors call a cause of “the race dynamic.” If the reward for being first evaporates within quarters, every lab is pushed to reach the next level sooner. That is the dynamic a pacing agreement would have to counteract, and why Anthropic says its numbers would shift if pacing happened.

Link 3: prices move under the pace metrics

Anthropic says compute is “among the most verifiable inputs” and so could be “a critical lever in a future pacing effort.” But a unit of compute buys very different amounts of work over time, and Epoch shows it does so at different rates in different domains: roughly 16–19× a year on math, about 3.6× on the coding benchmark. If safety workloads and capability workloads get cheaper at different rates, a constant compute share means a changing share of work.

Tool · What a flat share hides

Predict first
Safety share of compute (held flat)6%
Annual cost decline, safety workloads3.6×
Annual cost decline, capability workloads13×
Years2
What this shows

A thought experiment, not a claim about Anthropic’s workloads. The rates are Epoch’s measured spread across benchmarks (coding ~3.6×/yr, average ~13×/yr), and which rate applies to which workload is hypothetical. “Work” means compute divided by the cost of a fixed level of performance. Swap the two rates and the distortion runs the other way.

The lesson is not that safety is being starved. The rates could just as well run the other way, and Anthropic reports its own classifier overheads have moved in both directions. The lesson is that a pace metric denominated in compute needs a price deflator, and the deflator is different for different kinds of work. Epoch’s method, cost to reach a fixed level on a defined task, is exactly the kind of deflator such a metric would need. The two publications are two halves of one measurement programme, and neither is complete without the other.

Takeaways
  • Cheap thought makes agent workforces economic, and the AL4 share is what that looks like inside a lab
  • Fleeting premiums reward racing, which is the dynamic pacing must counter
  • Compute-denominated pace metrics need a work deflator, and Epoch’s method is one
7

Objections, and what each number licenses

Good-faith objections beyond the authors’ own caveats, and a tool for saying exactly what you can claim
By the end of this module you will
  • State the strongest objections to each measurement
  • Say precisely what each headline number does and does not let you claim
Objection 1 · Anthropic

The judge can drift with the thing it judges

The same model family gathers the evidence, built the task tree and judges the level. As Claude improves month to month, the research agent may also get better at finding AI involvement in the record, so ratings could rise partly from better detection. Blind human validation guards against this at one point in time; it does not by itself rule out drift over time.

The best reply: Anthropic proposes exactly the remedy: third-party verification, or other developers’ models as judges. The objection is an argument for doing that, not for dismissing the index.
Objection 2 · Anthropic

Self-measurement by an interested party

The measures arrive alongside Anthropic’s advocacy for pacing and for transparency rules. A lab has reasons to choose definitions that tell a particular story. On safety compute, it is the post itself that warns developers will be “tempted to draw the line generously.”

The best reply: the definitions are published, conservative by construction (dual-use work counts against safety), and paired with a commitment to embed independent evaluators with internal access. Incentive is a reason to verify, and the post asks to be verified.
Objection 3 · Epoch

Bounded benchmarks get cheap by construction

Once models comfortably exceed a benchmark’s ceiling, reaching any fixed score on it gets cheap quickly, because a small model can do it. The benchmarks closest to open-ended economic work, such as coding at 27.5% a quarter and FrontierMath Tier 4 at 26%, fall much more slowly. The headline may describe saturating tests better than frontier work.

The best reply: Epoch reports the slow benchmarks openly, and even the slowest rates are several times faster than any historical technology. The fair summary is “very fast, and slower where it matters most”.
Objection 4 · Epoch

List price is not cost

API prices can be strategic or subsidised, and the frontier premium Epoch describes is itself evidence that prices are not cost-based. Comparing a falling price for inference with falling costs for electricity or DNA sequencing mixes two things.

The best reply: for open-weight models Epoch used rented-hardware cost, which matched API pricing within 30% on five models it could check. And for users, price is what matters. The objection bites for claims about underlying technology, less for claims about affordability.

Tool · What does this number license?

Commit, then see why0 of 6 sorted

For each statement you might hear in a meeting, decide whether the publications support it, support it only with a qualifier, or do not support it.

“The cost of reaching a fixed benchmark score has been falling roughly an order of magnitude a year.”
Supported. This is the headline, stated at the right level of precision: ~13× a year, with a 42.9–58.0% quarterly range across reasonable averaging choices.
“Our AI coding costs will fall 13× next year.”
With a qualifier, and a big one. Coding was among the slowest-falling benchmarks (~3.6×/yr). Real users do not live on the frontier. And this is a historical rate, not a forecast.
“AI now does a quarter of Anthropic’s research on its own.”
Not supported. “Leads” means a human supervises and decides. Nothing was rated autonomous. And 26% is a share of person-time-weighted task categories, not of research output.
“Claude leads about a quarter of Anthropic’s AI R&D work.”
With a qualifier. It is Anthropic’s own figure, self-measured with a Claude judge, on a frozen July basket weighted by person-time, and it sits on the softest boundary of the scale.
“Only 6% of Anthropic’s effort goes into safety.”
Not supported. It is 6% of AI R&D compute in one week, conservatively classified, excluding safeguards classifiers. The authors stress that safety research is researcher-time heavy, not compute heavy.
“Every action by agents on Anthropic’s main internal platform passes through a monitor.”
Supported, as scoped: 100% coverage online and offline, on the most-used internal platform. What it does not tell you is how good the monitors are. Anthropic says labs should report how often known misbehaviour is caught.
Takeaways
  • Self-judging and self-reporting are arguments for third-party verification, which the post itself proposes
  • Price declines are fastest on saturating tests and slower where the work is most economically valuable
  • Say “leads”, not “does on its own”; say “compute share”, not “effort”