Two September 2026 publications try to put numbers on how fast AI is moving. Epoch AI measured how fast the price of a fixed level of AI performance is falling. Anthropic proposed how to measure, from inside a lab, how much of AI development AI is now doing itself. They look like different subjects. They are the same kind of instrument, with the same kinds of weakness, and one quietly depends on the other. This course teaches both, then puts them together.
Built on The plunging price of thought (Epoch AI, 22 Sep 2026) and Measurements for understanding the pace of AI development inside frontier labs (Anthropic, Sep 2026)
Both publications are careful about their own limits. Many of the sharpest caveats below are the authors’ own, and where that is so the course says so. Objections that go beyond what the authors concede are in Module 7 and are labelled as objections.
Short on time? Do Module 1 (the shared structure) and Module 6 (where the halves meet).
You cannot measure “thought” or “AI development” directly. Neither exists as a unit. So both publications do what statisticians do with inflation. They define a basket of things that can be measured, decide on a level to hold fixed or rate, choose weights to combine them, and choose a judge to score each item. The result is an index. Epoch says so explicitly: by averaging across benchmarks and performance levels, “we are effectively chaining price drops for various services. We are not tracking a conceptually unitary ‘price of thought.’”
| Design choice | Epoch: price of thought | Anthropic: R&D Automation Index |
|---|---|---|
| Basket | Five primary benchmarks: FrontierMath T1–3, OTIS Mock AIME, GPQA Diamond, Chess Puzzles, a secret Mystery Game | A frozen tree of 378 kinds of R&D task, built from July 2026 work records |
| Level held fixed / rated | A fixed accuracy score; measure the cheapest cost to reach it | An automation level, AL0–AL5, for each task |
| Weights | Averaged evenly over a grid of dates and accuracy levels | Person-time: each sampled person’s week split evenly across their tasks |
| Judge | The benchmark’s own answer key | An independent Claude judge, checked against staff ratings |
| Blind spot | Benchmarks are not useful work; nobody actually buys at the frontier | A frozen basket cannot see new kinds of work; Claude is judging Claude |
Index numbers have a small set of known failure modes: the basket drifts away from what matters, the weights encode a choice, the judge shares the errors of the thing judged, and a flat share can hide a changing quantity. Every one of them appears in these two publications. Learn to spot them once and you can read any AI-progress metric you will be shown next year.
Earlier estimates compared the per-token price of models that crossed a score threshold. Reasoning models broke that: they “can productively consume far more tokens, but as a result extract good performance from an underlying LLM that is smaller and cheaper to run per token.” A cheap-per-token model that thinks for ten times as long is not cheap. So Epoch measures the total cost to reach a score.
For each benchmark, every run gives a point: accuracy, release date, dollar cost. The cost frontier C(t, a) is the cheapest cost at which any model released by date t reaches accuracy a. The question is how fast that frontier falls while a is held fixed. It is the price of a fixed quality, the way economists track the price of computing power rather than the price of “a computer”.
Mapping a model’s whole cost-versus-accuracy curve would mean running it at many budgets. Instead Epoch follows a method from the federal Center for AI Standards and Innovation (CAISI). Take one long run’s transcript. For any per-question budget X, count only the questions answered correctly within X output tokens, and score the rest at the chance-guess rate (0.25 on four-way multiple choice). One run then yields a whole curve.
The obvious worry is that a model told its budget might do better than one silently cut off. Epoch tested this. Announcing a budget helped high-effort models at low budgets, but they did not beat low-effort variants run normally. So running the method across all thinking levels covers most of the achievable territory. Epoch also filled a gap it found in its own data: past benchmarking had chased peak scores, so it ran extra low-effort and small-model evaluations to avoid biasing the frontier.
A frontier is an extreme: one missing or extra run can move it for months. Sampling was “a bit ad hoc”, because Epoch mostly tested models near the frontier. And the grid is autocorrelated. So the authors write that they “abandon the formal paradigm of statistical inference”: they report reasonable but rough measurements, not confidence intervals. That candour is worth noticing. The right way to read their numbers is as a size, not a decimal.
“The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year.”Luke Emberson and David Roodman, Epoch AI
The worked example: OpenAI’s o3 (January 2025) reached 75% on GPQA Diamond for about 30 cents a question. GPT-5.6 Luna, just under 18 months later, matched it for $0.0004. The authors call it “a 725-fold drop in the price of thought in under 18 months”, like a $50,000 car falling to $69. By comparison, US residential electricity fell about 1.05× a year (1892–1973), lithium-ion batteries 1.16×, compute 1.51× and DNA sequencing 1.84×.
For three of the five main benchmarks, cost falls fastest right after a score first becomes state of the art. The authors’ explanation: the lab that gets there first can briefly charge a premium, then competitors, open and closed, catch up and the price plunges. Their implication is pointed: “supra-normal profits in LLM service provision may be fleeting for any given model,” which they call “surely both a consequence and a cause of the race dynamic.” Hold on to that line. It links price to pace in Module 6.
The post opens: “AI systems are getting more powerful, and they’re increasingly being used to build the next version of themselves.” Its stated aim is to show how close the world is to recursive self-improvement, “a model fully autonomously building its successor”. It is framed explicitly against Dario Amodei’s call to pace the frontier: Anthropic says it “would expect these numbers to shift if there were coordination on pacing.”
Anthropic’s own example makes the AL3/AL4 line concrete. A nightly data pipeline breaks. At AL3 the engineer brings the logs, stays tuned in, and Claude stops when something unexpected turns up. At AL4 the engineer hands over the alert. Claude finds the cause, fixes and tests it, handles surprises itself, reruns on a copy of the data and writes it up, then tags the engineer, who decides whether it ships. At AL5 nobody would have had to notice the failure at all.
The accompanying chart shows the AL4 share rising from under 1% in February 2026.
The 26% counts tasks rated AL4 rather than AL3. That boundary is exactly where Anthropic says ratings are softest: “There remains real room for disagreement on borderline cases, such as where exactly ‘AI collaborates’ ends and ‘AI leads’ begins.” Within one level 97% of the time is reassuring about the index overall. It says less about a number defined by a single level boundary. The 26% is a real signal, and its precision is loose.
Anthropic names the two obstacles to comparing labs: there is no common methodology, and the “judge” model “could make the same kinds of errors as the model it is checking.” Its proposed fixes are third-party verification or other developers’ models as judges. It also suggests these measures could become triggers, “like a fixed testing window before a new model is used for further AI R&D.”
About 30,000 agents did research and engineering work at any one time on Anthropic’s most-used internal platform in August 2026. The reason to measure: “individual agents rarely misbehave. But when there are millions or billions of agents operating in the economy, even rare events can happen regularly.”
| Online monitor | Offline monitor | |
|---|---|---|
| For | Fast, irreversible harms (for example copying weights out) | Slow, reversible harms (for example signs of misalignment) |
| Coverage | 100% of actions checked before execution, usually within seconds | 100% of actions ingested after the fact |
| Latency | Automated review before action; humans review blocked actions within a week | ~100,000 flagged transcripts a week triaged by classifiers; ~50 a week escalated to humans |
| Rate | 0.002% of over a billion August decisions blocked (about 1 in 47,000) | 1–2 transcripts in every thousand flagged |
Two design choices sit behind the numbers. Each agent has a persistent identity that survives model upgrades, so every action is auditable per agent. That lets an agent “treat what comes from another agent as a claim to check rather than a thought of its own.” And agents talk on an open, shared message system, cross-referenced with transcripts, so monitoring can follow communication between agents.
For one week in July 2026, Anthropic classified its compute. About 6% of compute going to AI R&D went to safety, and about 12% of compute going to AI-driven AI R&D. The estimates are “deliberately conservative”: work that helps capability as much as safety counts as R&D, and safeguards classifiers, “a separate, comparable amount of compute”, are excluded. The method: a Claude classifier reads each run’s metadata and code, on a 14% sample of the week’s ~10,000 runs, weighted toward the biggest. It agrees with human reviewers within a point or two.
The authors’ most important caveat is theirs, not a critic’s: “compute share measures only what is spent. A more efficient safety classifier, or a faster inference stack for production models, lowers the safety portion, but doesn’t mean we’re doing less safety work.” Module 6 puts a size on that sentence.
Thirty thousand agents running continuously is an inference bill. When the cost of a fixed level of performance falls about 13× a year, a workforce of agents at a given capability gets 13× cheaper to run each year, or 13× larger for the same money. Epoch measures the price; Anthropic measures the result: the share of R&D where AI leads went from under 1% to 26% in six months. Neither publication cites the other, but read together they describe one mechanism.
Epoch finds prices fall fastest just after a performance level becomes state of the art, so “supra-normal profits … may be fleeting for any given model”, which the authors call a cause of “the race dynamic.” If the reward for being first evaporates within quarters, every lab is pushed to reach the next level sooner. That is the dynamic a pacing agreement would have to counteract, and why Anthropic says its numbers would shift if pacing happened.
Anthropic says compute is “among the most verifiable inputs” and so could be “a critical lever in a future pacing effort.” But a unit of compute buys very different amounts of work over time, and Epoch shows it does so at different rates in different domains: roughly 16–19× a year on math, about 3.6× on the coding benchmark. If safety workloads and capability workloads get cheaper at different rates, a constant compute share means a changing share of work.
A thought experiment, not a claim about Anthropic’s workloads. The rates are Epoch’s measured spread across benchmarks (coding ~3.6×/yr, average ~13×/yr), and which rate applies to which workload is hypothetical. “Work” means compute divided by the cost of a fixed level of performance. Swap the two rates and the distortion runs the other way.
The lesson is not that safety is being starved. The rates could just as well run the other way, and Anthropic reports its own classifier overheads have moved in both directions. The lesson is that a pace metric denominated in compute needs a price deflator, and the deflator is different for different kinds of work. Epoch’s method, cost to reach a fixed level on a defined task, is exactly the kind of deflator such a metric would need. The two publications are two halves of one measurement programme, and neither is complete without the other.
The same model family gathers the evidence, built the task tree and judges the level. As Claude improves month to month, the research agent may also get better at finding AI involvement in the record, so ratings could rise partly from better detection. Blind human validation guards against this at one point in time; it does not by itself rule out drift over time.
The measures arrive alongside Anthropic’s advocacy for pacing and for transparency rules. A lab has reasons to choose definitions that tell a particular story. On safety compute, it is the post itself that warns developers will be “tempted to draw the line generously.”
Once models comfortably exceed a benchmark’s ceiling, reaching any fixed score on it gets cheap quickly, because a small model can do it. The benchmarks closest to open-ended economic work, such as coding at 27.5% a quarter and FrontierMath Tier 4 at 26%, fall much more slowly. The headline may describe saturating tests better than frontier work.
API prices can be strategic or subsidised, and the frontier premium Epoch describes is itself evidence that prices are not cost-based. Comparing a falling price for inference with falling costs for electricity or DNA sequencing mixes two things.
For each statement you might hear in a meeting, decide whether the publications support it, support it only with a qualifier, or do not support it.