The Attention Factory

OpenAI’s agentic pipeline gets read as “agents write the code now.” That is the least interesting thing in it. What the factory industrialises is deciding what a human should look at — context assembly, review routing, risk classification, alert de-duplication, incident triage. Every one of those is an attention router, and the system is bounded by the quality of its routers, not by the quality of its code generation.

9 modules
3 interactive tools
25 quiz questions
~55 min
15 Sep 2026

Built from reporting by Gergely Orosz at The Pragmatic Engineer and OpenAI’s own engineering documentation. Full sources at the end. The figure in Module 2 is original to this course.

Who this is for

An engineering leader deciding what, if anything, to take from OpenAI’s internal pipeline. The course is organised around one claim: the binding constraint in agentic delivery is human attention, and every interesting component in this system is a mechanism for allocating it. Read that way, the nine stages stop being a tour of features and become a set of design decisions you can accept, reject or copy one at a time.

Short on time? Module 4 for the prerequisite that actually gates this — CI throughput, not agent quality. Module 5 for the risk classifier, which is the crux, and why its errors make no sound. Module 7 for what to build first, in what order, and what the one measured example in this space actually did.

Course Modules

  1. A factory that routes attentionThe spine
  2. Nine stages, and who holds the penAnatomy
  3. Docs in the codeContext
  4. The hidden prerequisiteCI
  5. The risk classifier is the cruxThe crux
  6. The closing loopsFactory
  7. What you would need firstReadiness
  8. Three decisions, scoredInteractive
  9. CounterargumentsBoth sides
1

The reframe: a factory that routes attention

The code generation is the cheap part. The expensive part is deciding who looks
By the end of this module you will
  • Be able to state what the pipeline’s actual output is, other than software
  • Be able to name the attention routers in your own organisation today
  • Know why a router failure is structurally different from a code failure

Start from the step everyone skips

Step 3 of the nine — Codex actually writing the code — gets two sentences, opening with: “This part is trivial enough: Codex gets to work and makes a series of code changes until it reaches its goal, and then verifies that the software works as it should.”

Everything else in the pipeline is longer, more specific and more contested. That imbalance is the tell. If writing the code were the hard part, this would be a story about model capability. Instead it is about context assembly, load, review routing, risk classification, rollout supervision, alert de-duplication and incident triage — seven jobs, none of which is code generation, and every one of which is a decision about where a scarce resource gets spent.

The spine of this course

The factory does not automate writing software. It automates deciding what a human should look at.

Codex writing the change is throughput. The risk classifier deciding whether anyone reads it, the specialist reviewers deciding what is worth surfacing, the perf harness deciding which PRs warrant deeper evaluation, Perf Factory deciding which alerts are real, Sevbot deciding what is worth waking someone for — those are routers. They allocate the one input that did not get cheaper.

The quality of the system is bounded by the quality of its routers, not by the quality of its code generation. This is why “but our models aren’t as good as theirs” is the wrong objection, and why Module 4’s prerequisite — CI throughput — is a harder blocker than model quality for almost everyone.

Why this reframe earns its place over the obvious one

The common reading is: the human moved from writing code to defining outcomes and adjudicating risk. That is true. Venkat Venkataramani, who runs Applied Infra at OpenAI, describes engineers at OpenAI “are becoming more like product managers than traditional systems engineers,” and step 1 is a human defining a desired outcome where “judgment, prioritization, and taste are becoming more important.”

But it stops one step short, and the step it stops short of is the one that matters operationally. Saying the human now adjudicates risk invites the question “which risks?” — and the answer is that the human adjudicates the risks an agent decided to show them. Adjudication is downstream of routing. The routers are the control surface; the adjudication is what falls out of it.

Three things follow from the reframe that do not follow from the simpler one:

ConsequenceWhy
Router failures are silent by constructionIf a classifier wrongly marks a risky change low-risk, the observable outcome is nothing. No human looked, so no human knows they should have. Contrast a bad code change, which fails a test, fails a build, or pages someone. A mis-routed change produces no signal until production does. Module 5 is entirely about this
The routers are where an organisation’s judgement is stored“What counts as risky here” used to live in the heads of senior engineers and get applied ad hoc during review. A classifier makes it explicit, versionable, and auditable — which is genuinely better — and also makes it a single artifact that can be wrong everywhere at once
You already have routers, and they are probably not written downThe CODEOWNERS file. The “ping me on anything touching billing” Slack convention. The label that skips staging. The team norm that migrations get two reviewers. These are attention routers with no owner and no evaluation set. Replacing them with an agent is a change of implementation, not a change of kind — which is both the reassuring and the worrying way to look at it
The transferable question

If you take one thing from this course into your own organisation, take this: list your attention routers, and for each one name who owns it and how you would find out if it were wrong. Most teams can complete the first column, struggle with the second, and cannot answer the third at all. That gap is the honest measure of how far you are from anything resembling a factory — much more than your model choice is.

Takeaways
  • The reporting calls code generation “trivial enough”; everything else is routing
  • The spine: the factory automates deciding what a human should look at
  • Adjudication is downstream of routing — humans judge what an agent surfaced
  • Router failures are silent by construction; there is no failing test for “nobody looked”
  • Your routers already exist and mostly have no owner and no evaluation set
2

Nine stages, and who holds the pen

The anatomy, read as a profile of decision authority rather than a flow
By the end of this module you will
  • Know all nine stages and what each one actually does
  • Be able to say, for each, whether a human or an agent holds decision authority
  • See why the human’s two remaining footholds are at opposite ends of the pipeline

The usual way to draw this is left to right, as a flow. That drawing answers “what happens next?” The question this course cares about is “who decides?” — so here the nine stages are plotted against decision authority instead. Same system, different axis.

← scroll the figure →
HUMAN decides AGENT PROPOSES human decides AGENT decides 1 2 3 4 5 6 6 7 8 9 CONTESTED Perf Factory → re-enters at stage 3 Sevbot → re-enters at stage 3 defineoutcome gathercontext implement CI + perfharness specialistreview risk gate deploysupervision observeproduction incidentresponse Original figure for this course · stage content reported by The Pragmatic Engineer · authority read is this course’s own
Decision authority, not flow. The human holds the pen at exactly two points: the beginning, where the outcome is defined, and the end, where a mitigation is applied. Everything between is agent-held. Stage 6 is drawn as a fork because where the gate sits is genuinely unsettled — see Module 5. The two dashed green returns are what make this a factory rather than a pipeline: production feeds back into implementation without passing through a human first.

The nine stages, as reported

throughout this table; the pipeline is attributed in the article to Venkat Venkataramani, VP of Engineering, Applied Infra.

1
A human builder defines the desired outcome. An engineer or PM specifies the problem and the outcome. “Judgment, prioritization, and taste are becoming more important for this phase.” Engineers are described as becoming “more like product managers than traditional systems engineers.”
human
2
Codex gathers context. OpenAI “has moved all its documentation inside of the source code.” Codex also reaches Git and GitHub, Slack and Notion, internal data sources including Databricks and Datadog, and internal Codex skills — “some of which are maintained by OpenAI’s Codex implementation itself.” New engineers are told to ask Codex any onboarding question.
agent
3
Codex implements the change. In full: “This part is trivial enough: Codex gets to work and makes a series of code changes until it reaches its goal, and then verifies that the software works as it should.”
agent
4
Build, test, CI — and a perf harness. The agent builds, runs tests, fixes breakages, opens the PR, then “babysits the PR until it’s ‘green’, fixing any CI failures and automatically updating the PR.” New: a perf harness sends problematic PRs to the Synthetics A/B framework to evaluate performance implications.
agent
5
Agentic code review. “Instead of using one generic AI code reviewer, OpenAI spins off multiple agents, each with a ‘domain specialist’ configuration” — seen internally as “equivalent to having a human domain expert from each relevant infrastructure team review every change.” The coding agent then babysits the comments and updates the PR.
agent
6
Risk classification and the gate. High-risk changes “might invoke more AI code reviews, or mandate that a human reviews it after the AI agents finish.” Low-risk: “areas of the codebase can opt in to an agent that will auto-approve low risk PRs, removing human acceptance as a bottleneck.” Compliance input can also be routed automatically.
contested
7
Agentic deploy. “After a human approves a change to go to production, it is assigned its own agent” told to handhold it to a safe rollout. For a feature flag the agent finds the flag, understands the change, decides which signals indicate success and failure, builds its own monitoring dashboard, and watches. Building its own dashboard, rather than reading one someone else made, is the newest capability in the whole pipeline.
agent
8
Observe production. Agent-generated dashboards plus the internal observability stack — logs, metrics, traces, wide events. The change since the previous visit: engineers used to create dashboards per service; now agents do it “at the granularity of per-change deployment.”
agent
9
Perf Factory and Sevbot close the loop. Perf Factory sifts alerts and dashboards, de-duplicates signals, identifies real latency regressions, root-causes and proposes fixes. Sevbot collects incident context, “determines possible mitigations (but never executes any)”, answers questions in Slack; “an engineer can tell it to apply a specific mitigation.”
shared
Read the shape, not the stages

The human’s two footholds are intent at the front and production authority at the back, and they are the two places where being wrong is expensive in a way no test can catch. That is a defensible design, not an accident — and it is also a design with a specific vulnerability: everything between those two points is governed by stage 6, and stage 6 is the one stage whose position is genuinely unsettled.

Takeaways
  • Plotted by authority rather than flow, the shape is: human, then agent for seven stages, then shared
  • The human’s footholds are intent and production authority — the two no test catches
  • Stage 4 adds a perf harness routing problematic PRs to the Synthetics A/B framework
  • Stage 7’s agent chooses its own success signals before building its dashboard
  • Everything between the two footholds is governed by stage 6, the one stage whose position is unsettled
3

Docs in the code

A one-line detail in stage 2 that is doing more work than it looks like
By the end of this module you will
  • Know why relocating documentation into the repository changes what an agent can do
  • Know the failure mode OpenAI themselves hit doing a related thing
  • Be able to tell the good version of this from the version that becomes a liability

The claim

Step 2, in one sentence: “OpenAI has moved all its documentation inside of the source code, which makes it easier for agents to understand more of the code.”

Alongside it, Codex reaches Git and GitHub, Slack and Notion, internal data sources such as Databricks and Datadog, and internal skills. And: “Codex is so ‘plugged’ into OpenAI that new engineers are directed to ask Codex any questions they have during onboarding because it has a surprising amount of context.”

Why the relocation matters more than the access

The connectors are the part people notice and the least interesting part — Slack, Notion and Datadog integrations are a procurement exercise. The relocation is the part that changes the physics, for three reasons:

PropertyDocs in a wikiDocs in the repo
RetrievalA separate index the agent has to decide to query, with its own relevance model and its own failure modesArrives with the code the agent is already reading. No retrieval decision, no relevance model, no second system to be down
StalenessDecays silently. Nothing in the wiki knows the code changedChanges in the same diff, gets reviewed with the change, and shows up in blame. The doc and the code drift together or not at all
Ownership“Documentation” — usually nobody, often a team that does not own the codeThe repo’s owning team, by construction, because it is in their review

The second row is the real one. Every organisation has tried to fix stale documentation with process and failed. Moving it into the artifact that is already reviewed on every change is not a documentation initiative; it is a colocation that makes the existing review process do the work. That is why it is worth copying even if you never deploy an agent.

And the failure mode, which OpenAI hit themselves

Colocation is not the same as concatenation

The obvious way to implement “docs in the code” is the wrong one: put everything in one large instruction file at the repository root. That file crowds out the task and the code it was meant to support, and it rots into stale rules nobody maintains, because nothing forces a change to it when the system changes.

Note how narrowly the working version is scoped. OpenAI’s engineering guide recommends you “iterate on an AGENTS.md file that unlocks agentic loops like running tests and linters to receive feedback” — framing it as a small, functional file that enables a loop, not a manual. And in the Codex deepdive, AGENTS.md files are described narrowly: they “tell the agent how to navigate the codebase, which commands to run for testing, and how to follow the project’s standards.”

So the distinction to hold: docs in the code means documentation living next to the thing it documents. It does not mean one enormous instruction file at the repository root. The first scales with the codebase because it is distributed through it; the second is a single unconditional context cost paid on every task, and it is the one artifact in an agentic setup with no natural deletion process. If you copy one idea from this module, copy colocation, and be suspicious of any file that only ever grows.

The part that should give you pause

Internal skills are described as ones “some of which are maintained by OpenAI’s Codex implementation itself.” Read that against stage 2: the agent’s context is partly assembled from artifacts the agent maintains. That is not obviously wrong — it is how a compiler bootstraps — but a compiler has a test suite that catches a bad bootstrap, and it is not clear what plays that role here. It is the closest thing in the pipeline to a loop with no external check on it.

Takeaways
  • “All documentation inside the source code” is the highest-leverage line in stage 2
  • It works by removing a retrieval decision, tying staleness to the existing review, and conferring an owner
  • Colocation, not concatenation — a growing root instruction file is the failure mode
  • Some internal skills are maintained by the agent itself — a loop with no obvious external check
4

The hidden prerequisite

The throwaway line on the diagram is the most expensive item in the system
By the end of this module you will
  • Know the magnitude of the load change, in the words of the person responsible for absorbing it
  • Be able to explain why CI, not model quality, is what stops most organisations
  • Know the one part of OpenAI’s own delivery path where this has not been solved

The number

Venkat Venkataramani, VP of Engineering, Applied Infra:

“The number of pull requests (PRs) per engineer is growing like a hockey stick. Every part of the build-test-deploy pipeline is seeing dramatically more load. We’re talking about roughly a 10x increase in load on some systems. At most companies, that kind of growth might happen over two or three years. At OpenAI, we see it in about six months.

That level of acceleration exposes bottlenecks everywhere: version control has to handle far more code being written and pushed, CI/CD systems have to scale with it, and production release processes have to absorb a much higher rate of change. Every month, we wake up to a new set of infrastructure scaling challenges to solve.”

And the sentence that should end any “we’ll get there once we finish this project” plan: “Just when we think we’ve created enough capacity for the next phase of growth, the model unlocks another wave of capabilities, which creates a new set of bottlenecks somewhere else in the system.”

This is not unique to OpenAI, only faster there. GitHub’s own platform data shows pull requests opened rising roughly fivefold over three years, with PRs and commits nearly doubling from the end of 2025 alone. The slope is industry-wide; OpenAI is simply further along it.

Why this is the gate, and not the agents

Put the reframe from Module 1 next to this quote and the argument completes itself. The routers only work if the thing they route has throughput. Every stage downstream of implementation assumes CI can absorb the volume:

StageWhat it silently assumes about CI
Agent babysits the PR until greenThat a full run is cheap enough to do repeatedly per PR. The agent’s loop is fix → rerun → fix. If your suite takes 40 minutes, you have bought a 40-minute iteration cycle for a 30-second fix, and the agent will happily burn a day on it
Specialist review agents in parallelThat you can afford several full analyses per change. This one is cheap in wall-clock and expensive in tokens
Perf harness to an A/B frameworkThat a performance-evaluation environment exists at all, and can be driven programmatically on a subset of PRs. Most organisations do not have this and would not call its absence a CI problem
Per-change deploy supervisionThat release processes can absorb a much higher rate of change — not more code, more deployments. This is usually the first thing to break and the last thing anyone budgets for
The claim this module makes

Most organisations cannot run this pipeline because their CI would fall over, not because their agents are not good enough. Agent quality is a purchasing decision that improves on someone else’s roadmap. CI throughput is a capital project on yours. The diagram gives the first one seven boxes and the second one a parenthetical.

There is a diagnostic here that costs you nothing: take your median PR’s full CI wall-clock time and multiply it by ten. If the answer is a number your infrastructure could serve tomorrow, this pipeline is a process question for you. If it is not, everything downstream is hypothetical until you have spent that money.

The part of OpenAI’s own path where this is unsolved

Sulman Choudhry, Head of Engineering, ChatGPT, on native mobile — where every update passes Apple’s and Google’s review, taking hours or days:

“Code generation is getting dramatically faster, but getting that code into users’ hands on native mobile is not. For Codex in particular, where usage is heavily mobile-first, that gap is already becoming painful for us and users… We should be aiming for a world where shipping code on native mobile is as fast as shipping on the web. Getting there will probably require some creative rethinking of what we ship, when we ship it, and what can be activated remotely. Today, we’re nowhere close.”

This is the most useful paragraph in the whole picture for a sceptical reader, and no diagram of the factory has a box for it. An organisation with a comparable external gate — a regulator, a certification cycle, a customer-scheduled release window, a hardware dependency — should read it as the shape of their own ceiling. The factory accelerates the part of delivery you control, and then queues behind the part you do not. Throughput moves to the constraint; it does not remove it.

Takeaways
  • ~10x load on some systems in about six months, described as two-to-three years of growth elsewhere
  • “Every month, we wake up to a new set of infrastructure scaling challenges” — it does not complete
  • Most orgs are blocked by CI throughput, not agent quality; one is a capital project, the other a purchase
  • Free diagnostic: median PR CI wall-clock × 10
  • Native mobile is “nowhere close” — throughput moves to the constraint, it does not remove it
5

The risk classifier is the crux

Everything downstream rests on an agent correctly deciding what is low-risk — and its errors make no sound
By the end of this module you will
  • Know where the gate sits, and why that is still a live question
  • Be able to explain why a classifier error is undetectable by the system that made it
  • Have a concrete list of what would have to be true before you trusted one

How the gate works

High-risk changes go through stricter processes — more AI review passes, or a mandatory human reviewer after the agents finish. Low-risk changes take an easier path, and this is the sentence that matters: “areas of the codebase can opt in to an agent that will auto-approve low risk PRs, removing human acceptance as a bottleneck and improving velocity.” Compliance input can be routed the same way, automatically, to another agent or to a person. The motivation is volume: with the quantity of pull requests now being produced, humans cannot review all code without assistance.

Two things in that sentence are worth separating, because the second is usually dropped. The first is that an agent decides, per change, whether a human is needed. The second is that this is opt-in per area of the codebase — a human decision governs the scope of the automation even where the per-change decision does not. That is the difference between a policy someone owns and a switch someone flipped, and if you copy one thing from the design, copy the opt-in.

Where exactly the gate sits is a live question

Agentic deploy is described as starting “after a human approves a change to go to production”, and OpenAI’s own published guidance says engineers “delegate the initial code review to an agent, but own the final review and merge process”, with humans “responsible for judgment and final sign-off.” Set against auto-approval on opted-in areas, there is a real ambiguity about whether any human touches a low-risk change before it ships.

That ambiguity is worth carrying, because it is the exact question you will have to answer for yourself. Is your gate at merge, at deploy, or nowhere? Those are three different systems with three different failure profiles, and the loosest reading of this pipeline — humans removed entirely from the low-risk path — is the one least supported by anything anyone has published. Design for the gate you actually want rather than the one the diagram implies.

A useful contrast

Anthropic runs the same risk-based triage and has taken the opposite position on the gate. Jarred Sumner: a human merges even low-risk changes, with agent merging stated as a future goal rather than current practice.

Two frontier labs, the same triage concept, opposite answers to the one question that matters. That contrast is worth more than either lab’s practice alone, because it settles something: the position of the gate is a choice, not a consequence of capability. Both have the models. One moved the gate and one did not, and nothing about the models forced either decision. Anyone who tells you the technology has removed human merge is describing a preference as an inevitability.

The wider pattern matches. Across the industry there is far more talk about dropping human code review than there is evidence of teams actually doing it — and where it does happen, it is AI startups building extra layers underneath for safer production rollouts, not teams simply removing the reviewer.

How good is the review the gate is built on?

The gate sits downstream of the AI review, so the review’s quality matters to it. The one figure available is from the head of Codex: around nine out of ten comments from their bespoke review model point out valid issues, which he puts at equal to or slightly better than human reviewers. That model was trained specifically for review and tuned for signal over noise, which is the part worth copying — OpenAI’s own guidance is that generalised models “often nitpick and provide a low signal to noise ratio,” and that a reviewer “must be trained specifically to identify P0 and P1-level bugs.” Overly verbose review output gets ignored exactly like noisy lint warnings.

Worth holding next to any throughput argument: OpenAI declines to claim the review makes anything faster. “Code review doesn’t necessarily make the pull request process faster, especially if it finds meaningful bugs — but it does prevent defects and outages.” The review agents are a defect mechanism that happens to be parallelisable, not a velocity mechanism. If you adopt them expecting cycle time to drop, you have bought the wrong thing.

Why the error is silent

This is the part that should determine how you feel about the design, and it is a structural argument rather than a claim about anyone’s classifier.

Classifier saysRealityWhat the organisation observes
High riskActually high riskA human reviewed it. Working as intended
High riskActually low riskVISIBLE Wasted review time. Engineers complain. You will hear about this within a week
Low riskActually low riskFast merge. Working as intended
Low riskActually high riskSILENT Nothing. No human looked, so no human knows they should have. There is no failing test for “nobody reviewed this”. The only detector is production

The two error types have wildly asymmetric visibility, and the feedback you receive is almost entirely from the harmless one. That has a nasty second-order effect: the pressure on the classifier will be one-directional. Every week someone will complain that a trivial change was sent for human review. Nobody will complain about the change that was waved through, because the complaint would have to come from a person who never saw it. A classifier tuned by the feedback it actually receives drifts toward permissiveness, and the drift is invisible by the same mechanism that made the original error invisible.

This is not an argument against classifiers. It is an argument that a classifier without a deliberate counterweight will drift, and that the counterweight has to be manufactured, because the environment will not supply it.

What would have to be true before you trusted one

Five conditions. The first three are cheap; the fourth is the one nobody does; the fifth is the one that actually protects you.

ConditionWhy
1The risk taxonomy is written down and owned by a personIf “risky” is only in the model, you cannot review the policy, only its outputs. Duckbill’s version is a sentence: public API, auth, design system, non-additive schema changes, agent skills
2It is opt-in per area, not globalIt bounds blast radius to areas someone chose, and makes expansion a decision with a date on it
3Irreversibility is in the taxonomy, not just impactA change that is easy to revert and one that migrates data can have identical blast radius and completely different consequences for being wrong. Classify on how hard it is to undo, which is a property you can determine statically
4A sampled audit of the auto-approved pathThe counterweight. Pull a random sample of auto-approved changes each week and have a human review them after merge, recording how many should have been escalated. This is the only thing in the list that manufactures the missing signal, and I have not seen an organisation describe doing it
5A blameless path for “this should not have been auto-approved”Because the only other detector is an incident, and incidents make people defensive about the mechanism that caused them

Note that condition 4 is a specific instance of advice OpenAI gives generally in its own guide for choosing an AI review tool: “Curate examples of gold-standard PRs… Save this as an evaluation set to measure different tools,” and “Define how your team will measure whether reviews are high quality.” The gap is that an evaluation set measures the reviewer; nothing in the published guidance measures the router.

Takeaways
  • The auto-approve path is opt-in per area of the codebase, not a global setting
  • Decide whether your gate sits at merge, deploy, or nowhere — three systems, three failure profiles
  • Anthropic uses the same triage but a human still merges: the gate position is a choice, not a capability limit
  • False negatives are silent; false positives are loud — so tuning pressure is one-directional
  • Classify on irreversibility, not just impact; and run a sampled post-merge audit, which nobody describes doing
6

The closing loops

Perf Factory and Sevbot are what make it a factory rather than a pipeline
By the end of this module you will
  • Know why a loop back from production is categorically different from a longer pipeline
  • Know exactly what Sevbot does and the one thing it never does
  • Be able to tell OpenAI’s internal Sevbot from the sample app with a similar name

Pipeline versus factory

A pipeline is a sequence with an end. A factory has feedback: the output of the process re-enters it as input, without a person carrying it back. That is the whole distinction, and it is why the last two items on the diagram matter more than their size suggests.

Everything in stages 1–8 could be described as a very good pipeline. Stages 8 and 9 turn it into something else: production observations become new work, automatically, and the work re-enters at implementation. The system generates its own inputs.

Perf FactorySevbot
TriggerContinuous — alerts and dashboardsAn incident is detected; the bot “wakes up”
What it does “sift through alerts and dashboards, de-duplicate signals, identify real latency regressions, root-cause them and propose fixes”“Collects context about the incident”; “Determines possible mitigations (but never executes any)”; “Answers devs’ questions (it’s part of the Slack channel)”
Human’s roleThe fixes are proposed, so they re-enter the normal pipeline and are therefore subject to the stage 6 gate“An engineer can tell it to apply a specific mitigation”
Built onAgents“unsurprisingly, it’s also built on top of Codex”
Stated future —“OpenAI’s goal is to get to the point where Sevbot can take autonomous action when mitigating some outages. The dream is that no humans be woken up outside of their working hours… But as of now, oncall duty is not a thing of the past at the company.”

The most important word in the Sevbot description

“but never executes any”

The parenthetical is doing enormous work. Sevbot is a proposal engine with a human trigger, and autonomous mitigation is explicitly a goal rather than a state. Any summary of this system that says “an agent handles incidents” has dropped the only clause that determines whether it is a reasonable design.

Notice also the structural consistency with Module 2’s figure: the two places a human retains unambiguous authority are defining intent and taking irreversible production action. Sevbot is the second one. That is not a coincidence — it is the same principle applied at the other end of the pipeline, and it is the strongest evidence that the design has a coherent theory behind it rather than being automation applied wherever it fit.

A confusable worth knowing about

OpenAI publishes a developer showcase called “SRE Agent for Incident Response” whose sample app is literally named sev bot. It investigates incidents in Slack using GitHub and AWS, references past incidents, and proposes remediation. It is a runnable sample application built on the Agents API, not OpenAI’s internal system. Its own documentation says “Only a responder can approve a production change,” that the agent calls propose_rollback but cannot execute it, that “approval records a decision, not an executed rollback”, that “this app did not execute a rollback”, and that the first run uses bundled sample telemetry.

Two reasons to care. First, if you go looking for evidence about Sevbot you will find this, and it is easy to mistake a demo for a description of production. Second — and more usefully — the sample encodes the same constraint as the internal system: propose, never execute. When a vendor’s reference implementation and its internal practice agree on a boundary, that boundary is a considered position rather than a limitation of the current model.

What the loops buy, and what they cost

BuysCosts
De-duplication is the real product. “Sift through alerts… de-duplicate signals… identify real latency regressions”. Most alert estates are dominated by duplicates and noise; an agent that collapses them is doing the job a human oncall spends their first twenty minutes onA new router with the same silent-failure property as the classifier. A de-duplicator that wrongly merges two distinct regressions, or filters a real one as noise, produces no signal. Module 5’s whole argument applies here unchanged, and nobody describes measuring it
Work generation without a human courier. A latency regression becomes a proposed fix that re-enters at implementation. The organisation stops losing findings to “someone should file a ticket”The pipeline can now feed itself faster than it is inspected. Auto-generated work entering a pipeline with an auto-approving gate is the one composition in this system with no human in it anywhere, end to end — if the fix lands in an opted-in area
Incident context assembly, which is nearly pure gain. Collecting logs, commits and infrastructure changes is the tedious part of triage and is low-risk to automate because a human still decidesDependence concentrates. Separately in the same article: OpenAI is so dependent on Codex that in even a minor outage, colleagues’ messages reach the Codex team “at the same time as — or before — automated alerts.” An incident agent built on Codex has a correlated failure with the thing it exists to help debug

That last cell is the one I would raise in a design review. Building your incident-response agent on the same substrate as your development agent is efficient and creates a dependency that is exactly wrong in the tail: the scenario where you most need incident tooling is the scenario where the shared substrate may be the thing that is broken.

Takeaways
  • Loops, not stages, are what make it a factory: the system generates its own inputs
  • De-duplication is the real product of Perf Factory, and it is another silent-failure router
  • Sevbot never executes a mitigation; autonomous action is a stated goal and oncall still exists
  • OpenAI’s public sample app of the same name is a demo, and encodes the same propose-never-execute boundary
  • Auto-generated work entering an auto-approving gate is the one end-to-end human-free composition in the system
  • An incident agent on the same substrate as the dev agent has a correlated failure in the tail
7

What you would need first

Ordered so that nothing you build early is invalidated by something you learn late
By the end of this module you will
  • Have an ordered list of prerequisites, with the reason each one sits where it does
  • Know which of them the only measured example in this material actually built
  • Know which stages are worth copying early and which are worth copying last or never

The ordering principle

Do first the things whose absence would invalidate the work that follows. In practice this means the unglamorous items come first, because the glamorous ones depend on them and not the reverse. Every step below is a precondition for the ones under it, and the pipeline in Module 2 is the last row, not the first.

#PrerequisiteWhy here
0A verification layer you would bet onEverything in the factory is a loop that terminates when a check passes. If your tests do not fail when the code is wrong, every agent in the pipeline is confidently producing garbage at speed and the routers are routing noise. OpenAI’s own guide puts it the same way: “defining high quality tests is often the first step to allowing an agent to build a feature.” This is not step one because it is virtuous; it is step one because nothing downstream means anything without it
1CI throughput sized for ten times the current PR rateModule 4. A capital project, on your roadmap, that improves nothing visible on the day it lands. Skipping it does not make the pipeline slower; it makes the pipeline impossible, because the agent’s inner loop is fix-rerun-fix
2Documentation colocated with codeModule 3. Cheap, independently valuable, and the thing that makes every later stage’s context better. Also the one item here you would do even if you abandoned the whole plan
3A written risk taxonomy with a named ownerModule 5. Note this comes before any classifier: you cannot evaluate a classifier against a policy that does not exist, and writing the policy down is most of the value. Duckbill’s fits in a sentence and was enforced with a shell script adding a GitHub label — no model involved
4An evaluation set for review quality, before adopting a reviewerStraight from OpenAI’s guide: “Curate examples of gold-standard PRs… Save this as an evaluation set to measure different tools”, and “Define how your team will measure whether reviews are high quality.” Adopting first and evaluating later means you never find out, because by then the baseline is gone
5Post-merge sampled audit of anything auto-approvedModule 5, condition 4. Build this with the auto-approval, never after — if it arrives later you have an unmeasured period you can never reconstruct, and the drift will already have happened
6Feature flags and a rollback you have actually usedAgentic deploy supervises a rollout; it does not create the ability to roll one back. If your current recovery story is “revert and redeploy, about forty minutes,” then per-change agent supervision is watching a fire it cannot put out
7Alert hygiene sufficient for de-duplication to be meaningfulModule 6. An agent that de-duplicates a noisy alert estate makes the noise tolerable, which removes the pressure to fix it. Doing this before the estate is sane buys you a permanent, automated excuse
8Only now: the pipelineAnd even then, incrementally, by area, in the order in Module 2 — context assembly first, auto-approval last

What the only measured example actually did

Duckbill, a fifteen-person company, via its cofounder and CEO Mike Julian. Their trigger was recognisably normal: “we found ourselves with 60 open PRs for a team of five… we all had the sudden realization we were looking at two days of just code review.”

What they built before removing review, in their own description:

  • A risk-based system: human review required if a change touched the public API/MCP, auth, the design system, non-additive database schema changes, or agent skills — “We then enforced that with a shell script to add a GitHub label”
  • Guardrails: “We enabled nearly every rule in ruff/prettier/eslint/ty, and we improved our unit test coverage to a floor of 85%”
  • Their own assessment of the cost: “Improving guardrails was pretty easy, just expensive in tokens and attention”

And the result: PRs merged 353 → 684 (80/wk → 154/wk, +94%); merged within an hour 28% → 45%; human-reviewed median merge time 26 hours, versus 1 hour without.

Why this example is more useful to you than OpenAI’s

Three reasons. It is a size you can reason about. It has an actual before and after. And its risk gate is a shell script and a label — no classifier, no model, no agent. The taxonomy did the work; the mechanism enforcing it was deliberately boring.

That is the most transferable finding in this entire course. The valuable half of the risk gate is the taxonomy, and the taxonomy is free. You can have most of the benefit described in the diagram — humans spending their review attention where blast radius is real — with a rule file and a linter, and you can have it this quarter, with none of the silent-failure exposure of a learned classifier. Start there. Move to a classifier when the deterministic rule genuinely cannot express the distinction you need, and when you have the audit from step 5 running.

What to copy early, late, or never

StageVerdictReason
Docs colocated with codeCOPY NOWValuable with or without agents; improves every later stage
Agent babysits PR to greenCOPY NOWBounded, reversible, obviously checkable — CI either passes or it does not. Gated only by step 1
AI review as an added laneCOPY NOWAdditive. It comments; humans still decide. Do step 4 first so you can tell whether it is any good
Written risk taxonomy + deterministic enforcementCOPY NOWThe Duckbill finding. Free, and it is the valuable half
Incident context assembly (propose, never execute)COPY NOWTedious work, human keeps authority, and the boundary is one both OpenAI’s internal and reference implementations hold
Specialist review agentsLATERA context-engineering exercise whose value is unproven even in the source. Real cost, plausible benefit
Agentic deploy with self-authored dashboardsLATERRequires step 6 and a mature observability stack. The agent choosing its own success signals is a bigger delegation than it reads as
Learned risk classifier auto-approving mergesLASTSilent failure mode, no accuracy evidence in any source, and both frontier labs disagree about whether to do it. Last, opt-in per area, with the audit running
Autonomous incident mitigationNOT YETNobody has it. It is a stated goal at the company that would most plausibly have built it already
Takeaways
  • Order by what would invalidate later work: verification, then CI, then docs, then taxonomy
  • The risk taxonomy comes before any classifier — you cannot evaluate against a policy that doesn’t exist
  • Build the audit with the auto-approval, never after
  • Duckbill’s gate was a shell script and a label; the taxonomy is the valuable half and it is free
  • Copy now: colocated docs, PR-to-green, AI review as a lane, written taxonomy, incident context assembly
  • Copy last or never: learned auto-approval, autonomous mitigation
8

Three decisions, scored

Who holds the pen, whether a failure makes a sound, and what you would build first
By the end of this module you will
  • Have assigned decision authority across the pipeline without looking at the figure
  • Have worked out which failures in this system produce a signal and which produce silence
  • Have ordered a real readiness programme and found out where you put the glamorous parts

Nine situations. For each, say who holds decision authority in the pipeline as it actually runs — not as you would design it.

Tool 1

Who holds the pen?

0 / 9
Result

Tool 2 — will you find out?

Ten things go wrong. For each, name the detector: does an automated check catch it, does a human notice, or does nothing happen at all until production tells you? This is the tool worth doing twice. The pattern in the answers is the whole argument of the course.

Tool 2

Which failures make a sound?

0 / 10
Result

Tool 3 — what would you build first?

Eight candidate projects. For each: now (cheap, additive, and it unlocks later work), later (real value, but it depends on something you do not have yet), or not yet (nobody can yet show it working).

Tool 3

Now, later, or not yet

0 / 8
Result

Takeaways
  • Authority in the reported system: human at intent, human at irreversible action, agent between
  • Five of the ten failures produce no signal at all, and every one of those five is a router failure
  • Readiness ordering puts the boring capital work first and the classifier last
9

Counterarguments

Marked as objections. The last one is that this works and generalises, and it is not a straw man
By the end of this module you will
  • Know the strongest cases against taking any of this seriously
  • Know the strongest case for taking all of it seriously
  • Know which ones I think are right
Objection 1 — and I think this one lands hardest

OpenAI is the least representative organisation that could possibly have produced this

Start with the input that changes everything downstream: every engineer, researcher, and finance and marketing colleague at OpenAI works with an unlimited token budget. No usage cap on LLMs, the same as at Anthropic. Parallel specialist reviewers on every change, a supervising agent per deployment, nightly passes over the whole codebase — all of that is rational when inference is free at the margin and none of it is obviously rational when it is not.

And the tooling is not tooling you can buy. OpenAI’s internal Codex is substantially more advanced than the product, because it is plugged into pretty much every internal system. Add that they train the model, see its failure modes before anyone else, and can change the model in response to what the pipeline needs — and that the engineers are a heavily selected population. Several of these inputs cannot be purchased at any price.

What bears on it: This is correct and it is the most important objection here. But notice what it does and does not undermine. It undermines the economics — the parallel specialist reviewers, the per-change deploy agents, the always-running nightly passes all assume inference is free at the margin, and for you it is not. It does not undermine the structure: that the binding constraint is human attention, that the routers are the control surface, and that router errors are silent, are all properties of the design rather than of the budget. Read the pipeline for its shape and discount every cost assumption to zero confidence. Also note the objection cuts the other way on one point: an organisation with a finite token budget has a natural forcing function toward the cheapest mechanism that does the job, which OpenAI does not.
Objection 2

A diagram of an internal system is a recruiting and marketing artifact as much as a description

OpenAI sells Codex. A story about OpenAI running on Codex is the best available advertisement for Codex, and a recruiting pitch to exactly the engineers they want. The people describing the pipeline are leaders whose standing is tied to it working. No sceptic speaks; nobody describes a change that went wrong; there is no failure story anywhere in nine stages. And OpenAI’s engineering guide closes with “reach out to OpenAI. We’re here to help you turn coding agents into real leverage” — explicitly a sales document.

What bears on it: Right about the incentives, too strong as a conclusion. Three things cut against it. Plenty of what is on record is material no marketing department would approve: native mobile “nowhere close”, oncall not a thing of the past, and “every month, we wake up to a new set of infrastructure scaling challenges.” OpenAI’s own guide declines to claim AI review makes pull requests faster, which is precisely the claim a sales document would lead with. And a selection effect is not a fabrication — the useful response is to read the design for its shape rather than to dismiss it. The residue that survives all three: there are no failure stories, and a system with no failure stories has not been described completely. When you copy from it, budget for the failures nobody mentioned.
Objection 3

Removing humans from the low-risk path optimises throughput on exactly the changes where throughput matters least

By construction, low-risk changes are the ones where being wrong is cheap. They are also usually the small ones. So the auto-approve gate buys speed on the cheap half of the work while carrying the entire tail risk of misclassification — because every catastrophic outcome of this design lives in the case where something believed low-risk was not. The expected value is a modest, diffuse throughput gain against a rare, concentrated, and structurally undetectable loss. That is a bad shape for a bet, and it is the shape you get whenever you optimise the common case by removing the check that catches the uncommon one.

What bears on it: I think this is the most underrated objection in the set and I hold it partly. What it gets right: the payoff really is asymmetric, and Module 5 shows the loss side is undetectable, which is worse than merely rare. What it misses is the second-order effect, which is the actual argument for the design. The benefit is not that low-risk changes merge faster; it is that reviewer attention is withdrawn from work where it was doing nothing and is available for work where it does something. If reviewing 200 trivial PRs a week is why nobody reads the migration carefully, then removing the 200 improves the review of the migration. Whether that is what actually happens is an empirical question nobody in these sources has answered — and if the freed attention goes to more tickets rather than to deeper review, the objection wins outright. That is a measurement any adopting organisation could take and, as far as I can tell, none has.
Objection 4

Analysing the pipeline this carefully is itself a way of not deciding

An organisation can study attention routing, failure asymmetries and readiness ordering for two quarters and ship none of it, while the teams actually getting value run the crude version: turn on an AI reviewer, read the comments for a fortnight, keep it or kill it. Careful analysis is more comfortable than that, and from the outside it is indistinguishable from delay — especially to a leadership team that would rather not commit.

What bears on it: Largely fair, and it is why Module 7 is biased toward action: five of its nine verdicts are “copy now”, and the cheapest of them cost nothing. The analysis earns its place at two moments and not many others — when someone puts this pipeline in front of leadership as a plan, and when you are deciding whether auto-approval is the next thing you build. Everywhere else, the crude version is the right move. If the choice is between another week of reading and turning on an AI reviewer behind a written risk taxonomy, turn on the reviewer.
Objection 5 — the case for the pipeline, and it is not a straw man

This works, and it generalises, and the sceptical reading will age badly

The honest case: every part of this pipeline is a thing that was impossible two years ago and is now routine at one company, and the pattern of such things is that they diffuse. The prerequisites in Module 7 are real today and are exactly the kind of thing that gets commoditised — CI capacity is a purchasing decision the moment someone sells agent-scale CI, and they will. The specialist-reviewer mechanism is plausible on its face and cheap to test. The risk taxonomy is free, as Module 7 itself argues. And the silent-failure argument, while structurally sound, proves too much: human code review has exactly the same asymmetry — a reviewer who skims and approves also produces no signal — and we have run the industry on it for forty years without demanding a sampled audit. Holding agents to a standard we never applied to ourselves is not rigour, it is status-quo bias with a ledger.

What bears on it: The last sentence is the strongest thing in this module and I do not have a clean answer to it. Three responses, and only the third is any good. The weak one: scale — a skimming reviewer is one person on one PR, a miscalibrated classifier is one policy on every PR, so the variance is different even if the bias is not. The weaker one: we did in fact build counterweights to human review, they were just informal — the colleague who wanders over, the reviewer who happens to know that subsystem — and those disappear rather than transfer. The one I actually believe: the objection is right that the standard is asymmetric, and the correct conclusion is not to relax the standard for agents but to notice we should have had the audit all along. A sampled post-merge review is a good idea for human-reviewed changes too, and almost nobody does that either. If this pipeline’s main legacy turns out to be that it forced organisations to write down what “risky” means and to measure whether their routing was right, that is a real gain independent of whether any agent was ever involved.

Three arguments that do not survive the detail

ArgumentWhy it fails
“OpenAI has removed humans from software delivery.”The article says agentic deploy begins after a human approves the change, that Sevbot never executes a mitigation, and that oncall duty is not a thing of the past. OpenAI’s own guide says engineers own the final review and merge. Module 5.
“This is just a normal CI pipeline with AI buzzwords on it.”The two feedback loops are a category difference: production observations become new work without a human courier, and the work re-enters at implementation. Module 6.
“We can’t do this because our models aren’t as good as theirs.”Model access is a purchase. What stops most organisations is CI throughput, a verification layer worth betting on, and a written risk taxonomy — and the measured example in this material is a fifteen-person company whose risk gate was a shell script. Modules 5 and 9.
If you do one thing after this course

Write down what “risky” means in your codebase, in one paragraph, and have one person own it. Enforce it with the dumbest mechanism that works. Then look at what your engineers currently spend review attention on and ask whether that paragraph is where it is going. You will have implemented the valuable half of the most contested box in the diagram, this week, with no agent involved — and you will have a policy you can actually evaluate a classifier against, on the day you decide you want one.

Takeaways
  • Objection 1 is right: discount every cost assumption to zero, keep the structure
  • Objection 2 is partly right; the residue is that there are no failure stories
  • Objection 3 is underrated: the payoff is diffuse gain against concentrated undetectable loss
  • Objection 5 lands: human review has the same asymmetry, and we never audited that either
  • Do not say “OpenAI removed humans”, “it’s just CI”, or “our models aren’t good enough”
  • This week: write down what “risky” means, name an owner, enforce it with the dumbest mechanism that works