OpenAI’s agentic pipeline gets read as “agents write the code now.” That is the least interesting thing in it. What the factory industrialises is deciding what a human should look at — context assembly, review routing, risk classification, alert de-duplication, incident triage. Every one of those is an attention router, and the system is bounded by the quality of its routers, not by the quality of its code generation.
Built from reporting by Gergely Orosz at The Pragmatic Engineer and OpenAI’s own engineering documentation. Full sources at the end. The figure in Module 2 is original to this course.
An engineering leader deciding what, if anything, to take from OpenAI’s internal pipeline. The course is organised around one claim: the binding constraint in agentic delivery is human attention, and every interesting component in this system is a mechanism for allocating it. Read that way, the nine stages stop being a tour of features and become a set of design decisions you can accept, reject or copy one at a time.
Short on time? Module 4 for the prerequisite that actually gates this — CI throughput, not agent quality. Module 5 for the risk classifier, which is the crux, and why its errors make no sound. Module 7 for what to build first, in what order, and what the one measured example in this space actually did.
Step 3 of the nine — Codex actually writing the code — gets two sentences, opening with: “This part is trivial enough: Codex gets to work and makes a series of code changes until it reaches its goal, and then verifies that the software works as it should.”
Everything else in the pipeline is longer, more specific and more contested. That imbalance is the tell. If writing the code were the hard part, this would be a story about model capability. Instead it is about context assembly, load, review routing, risk classification, rollout supervision, alert de-duplication and incident triage — seven jobs, none of which is code generation, and every one of which is a decision about where a scarce resource gets spent.
The factory does not automate writing software. It automates deciding what a human should look at.
Codex writing the change is throughput. The risk classifier deciding whether anyone reads it, the specialist reviewers deciding what is worth surfacing, the perf harness deciding which PRs warrant deeper evaluation, Perf Factory deciding which alerts are real, Sevbot deciding what is worth waking someone for — those are routers. They allocate the one input that did not get cheaper.
The quality of the system is bounded by the quality of its routers, not by the quality of its code generation. This is why “but our models aren’t as good as theirs” is the wrong objection, and why Module 4’s prerequisite — CI throughput — is a harder blocker than model quality for almost everyone.
The common reading is: the human moved from writing code to defining outcomes and adjudicating risk. That is true. Venkat Venkataramani, who runs Applied Infra at OpenAI, describes engineers at OpenAI “are becoming more like product managers than traditional systems engineers,” and step 1 is a human defining a desired outcome where “judgment, prioritization, and taste are becoming more important.”
But it stops one step short, and the step it stops short of is the one that matters operationally. Saying the human now adjudicates risk invites the question “which risks?” — and the answer is that the human adjudicates the risks an agent decided to show them. Adjudication is downstream of routing. The routers are the control surface; the adjudication is what falls out of it.
Three things follow from the reframe that do not follow from the simpler one:
| Consequence | Why |
|---|---|
| Router failures are silent by construction | If a classifier wrongly marks a risky change low-risk, the observable outcome is nothing. No human looked, so no human knows they should have. Contrast a bad code change, which fails a test, fails a build, or pages someone. A mis-routed change produces no signal until production does. Module 5 is entirely about this |
| The routers are where an organisation’s judgement is stored | “What counts as risky here” used to live in the heads of senior engineers and get applied ad hoc during review. A classifier makes it explicit, versionable, and auditable — which is genuinely better — and also makes it a single artifact that can be wrong everywhere at once |
| You already have routers, and they are probably not written down | The CODEOWNERS file. The “ping me on anything touching billing” Slack convention. The label that skips staging. The team norm that migrations get two reviewers. These are attention routers with no owner and no evaluation set. Replacing them with an agent is a change of implementation, not a change of kind — which is both the reassuring and the worrying way to look at it |
If you take one thing from this course into your own organisation, take this: list your attention routers, and for each one name who owns it and how you would find out if it were wrong. Most teams can complete the first column, struggle with the second, and cannot answer the third at all. That gap is the honest measure of how far you are from anything resembling a factory — much more than your model choice is.
The usual way to draw this is left to right, as a flow. That drawing answers “what happens next?” The question this course cares about is “who decides?” — so here the nine stages are plotted against decision authority instead. Same system, different axis.
throughout this table; the pipeline is attributed in the article to Venkat Venkataramani, VP of Engineering, Applied Infra.
The human’s two footholds are intent at the front and production authority at the back, and they are the two places where being wrong is expensive in a way no test can catch. That is a defensible design, not an accident — and it is also a design with a specific vulnerability: everything between those two points is governed by stage 6, and stage 6 is the one stage whose position is genuinely unsettled.
Step 2, in one sentence: “OpenAI has moved all its documentation inside of the source code, which makes it easier for agents to understand more of the code.”
Alongside it, Codex reaches Git and GitHub, Slack and Notion, internal data sources such as Databricks and Datadog, and internal skills. And: “Codex is so ‘plugged’ into OpenAI that new engineers are directed to ask Codex any questions they have during onboarding because it has a surprising amount of context.”
The connectors are the part people notice and the least interesting part — Slack, Notion and Datadog integrations are a procurement exercise. The relocation is the part that changes the physics, for three reasons:
| Property | Docs in a wiki | Docs in the repo |
|---|---|---|
| Retrieval | A separate index the agent has to decide to query, with its own relevance model and its own failure modes | Arrives with the code the agent is already reading. No retrieval decision, no relevance model, no second system to be down |
| Staleness | Decays silently. Nothing in the wiki knows the code changed | Changes in the same diff, gets reviewed with the change, and shows up in blame. The doc and the code drift together or not at all |
| Ownership | “Documentation” — usually nobody, often a team that does not own the code | The repo’s owning team, by construction, because it is in their review |
The second row is the real one. Every organisation has tried to fix stale documentation with process and failed. Moving it into the artifact that is already reviewed on every change is not a documentation initiative; it is a colocation that makes the existing review process do the work. That is why it is worth copying even if you never deploy an agent.
The obvious way to implement “docs in the code” is the wrong one: put everything in one large instruction file at the repository root. That file crowds out the task and the code it was meant to support, and it rots into stale rules nobody maintains, because nothing forces a change to it when the system changes.
Note how narrowly the working version is scoped. OpenAI’s engineering guide recommends you
“iterate on an AGENTS.md file that unlocks agentic loops like running tests and linters to receive
feedback” — framing it as a small, functional file that enables a loop, not a manual.
And in the Codex deepdive, AGENTS.md files are
described narrowly: they “tell the agent how to navigate the codebase, which commands to run for
testing, and how to follow the project’s standards.”
So the distinction to hold: docs in the code means documentation living next to the thing it documents. It does not mean one enormous instruction file at the repository root. The first scales with the codebase because it is distributed through it; the second is a single unconditional context cost paid on every task, and it is the one artifact in an agentic setup with no natural deletion process. If you copy one idea from this module, copy colocation, and be suspicious of any file that only ever grows.
Internal skills are described as ones “some of which are maintained by OpenAI’s Codex implementation itself.” Read that against stage 2: the agent’s context is partly assembled from artifacts the agent maintains. That is not obviously wrong — it is how a compiler bootstraps — but a compiler has a test suite that catches a bad bootstrap, and it is not clear what plays that role here. It is the closest thing in the pipeline to a loop with no external check on it.
Venkat Venkataramani, VP of Engineering, Applied Infra:
“The number of pull requests (PRs) per engineer is growing like a hockey stick. Every part of the build-test-deploy pipeline is seeing dramatically more load. We’re talking about roughly a 10x increase in load on some systems. At most companies, that kind of growth might happen over two or three years. At OpenAI, we see it in about six months.
That level of acceleration exposes bottlenecks everywhere: version control has to handle far more code being written and pushed, CI/CD systems have to scale with it, and production release processes have to absorb a much higher rate of change. Every month, we wake up to a new set of infrastructure scaling challenges to solve.”
And the sentence that should end any “we’ll get there once we finish this project” plan: “Just when we think we’ve created enough capacity for the next phase of growth, the model unlocks another wave of capabilities, which creates a new set of bottlenecks somewhere else in the system.”
This is not unique to OpenAI, only faster there. GitHub’s own platform data shows pull requests opened rising roughly fivefold over three years, with PRs and commits nearly doubling from the end of 2025 alone. The slope is industry-wide; OpenAI is simply further along it.
Put the reframe from Module 1 next to this quote and the argument completes itself. The routers only work if the thing they route has throughput. Every stage downstream of implementation assumes CI can absorb the volume:
| Stage | What it silently assumes about CI |
|---|---|
| Agent babysits the PR until green | That a full run is cheap enough to do repeatedly per PR. The agent’s loop is fix → rerun → fix. If your suite takes 40 minutes, you have bought a 40-minute iteration cycle for a 30-second fix, and the agent will happily burn a day on it |
| Specialist review agents in parallel | That you can afford several full analyses per change. This one is cheap in wall-clock and expensive in tokens |
| Perf harness to an A/B framework | That a performance-evaluation environment exists at all, and can be driven programmatically on a subset of PRs. Most organisations do not have this and would not call its absence a CI problem |
| Per-change deploy supervision | That release processes can absorb a much higher rate of change — not more code, more deployments. This is usually the first thing to break and the last thing anyone budgets for |
Most organisations cannot run this pipeline because their CI would fall over, not because their agents are not good enough. Agent quality is a purchasing decision that improves on someone else’s roadmap. CI throughput is a capital project on yours. The diagram gives the first one seven boxes and the second one a parenthetical.
There is a diagnostic here that costs you nothing: take your median PR’s full CI wall-clock time and multiply it by ten. If the answer is a number your infrastructure could serve tomorrow, this pipeline is a process question for you. If it is not, everything downstream is hypothetical until you have spent that money.
Sulman Choudhry, Head of Engineering, ChatGPT, on native mobile — where every update passes Apple’s and Google’s review, taking hours or days:
“Code generation is getting dramatically faster, but getting that code into users’ hands on native mobile is not. For Codex in particular, where usage is heavily mobile-first, that gap is already becoming painful for us and users… We should be aiming for a world where shipping code on native mobile is as fast as shipping on the web. Getting there will probably require some creative rethinking of what we ship, when we ship it, and what can be activated remotely. Today, we’re nowhere close.”
This is the most useful paragraph in the whole picture for a sceptical reader, and no diagram of the factory has a box for it. An organisation with a comparable external gate — a regulator, a certification cycle, a customer-scheduled release window, a hardware dependency — should read it as the shape of their own ceiling. The factory accelerates the part of delivery you control, and then queues behind the part you do not. Throughput moves to the constraint; it does not remove it.
High-risk changes go through stricter processes — more AI review passes, or a mandatory human reviewer after the agents finish. Low-risk changes take an easier path, and this is the sentence that matters: “areas of the codebase can opt in to an agent that will auto-approve low risk PRs, removing human acceptance as a bottleneck and improving velocity.” Compliance input can be routed the same way, automatically, to another agent or to a person. The motivation is volume: with the quantity of pull requests now being produced, humans cannot review all code without assistance.
Two things in that sentence are worth separating, because the second is usually dropped. The first is that an agent decides, per change, whether a human is needed. The second is that this is opt-in per area of the codebase — a human decision governs the scope of the automation even where the per-change decision does not. That is the difference between a policy someone owns and a switch someone flipped, and if you copy one thing from the design, copy the opt-in.
Agentic deploy is described as starting “after a human approves a change to go to production”, and OpenAI’s own published guidance says engineers “delegate the initial code review to an agent, but own the final review and merge process”, with humans “responsible for judgment and final sign-off.” Set against auto-approval on opted-in areas, there is a real ambiguity about whether any human touches a low-risk change before it ships.
That ambiguity is worth carrying, because it is the exact question you will have to answer for yourself. Is your gate at merge, at deploy, or nowhere? Those are three different systems with three different failure profiles, and the loosest reading of this pipeline — humans removed entirely from the low-risk path — is the one least supported by anything anyone has published. Design for the gate you actually want rather than the one the diagram implies.
Anthropic runs the same risk-based triage and has taken the opposite position on the gate. Jarred Sumner: a human merges even low-risk changes, with agent merging stated as a future goal rather than current practice.
Two frontier labs, the same triage concept, opposite answers to the one question that matters. That contrast is worth more than either lab’s practice alone, because it settles something: the position of the gate is a choice, not a consequence of capability. Both have the models. One moved the gate and one did not, and nothing about the models forced either decision. Anyone who tells you the technology has removed human merge is describing a preference as an inevitability.
The wider pattern matches. Across the industry there is far more talk about dropping human code review than there is evidence of teams actually doing it — and where it does happen, it is AI startups building extra layers underneath for safer production rollouts, not teams simply removing the reviewer.
The gate sits downstream of the AI review, so the review’s quality matters to it. The one figure available is from the head of Codex: around nine out of ten comments from their bespoke review model point out valid issues, which he puts at equal to or slightly better than human reviewers. That model was trained specifically for review and tuned for signal over noise, which is the part worth copying — OpenAI’s own guidance is that generalised models “often nitpick and provide a low signal to noise ratio,” and that a reviewer “must be trained specifically to identify P0 and P1-level bugs.” Overly verbose review output gets ignored exactly like noisy lint warnings.
Worth holding next to any throughput argument: OpenAI declines to claim the review makes anything faster. “Code review doesn’t necessarily make the pull request process faster, especially if it finds meaningful bugs — but it does prevent defects and outages.” The review agents are a defect mechanism that happens to be parallelisable, not a velocity mechanism. If you adopt them expecting cycle time to drop, you have bought the wrong thing.
This is the part that should determine how you feel about the design, and it is a structural argument rather than a claim about anyone’s classifier.
| Classifier says | Reality | What the organisation observes |
|---|---|---|
| High risk | Actually high risk | A human reviewed it. Working as intended |
| High risk | Actually low risk | VISIBLE Wasted review time. Engineers complain. You will hear about this within a week |
| Low risk | Actually low risk | Fast merge. Working as intended |
| Low risk | Actually high risk | SILENT Nothing. No human looked, so no human knows they should have. There is no failing test for “nobody reviewed this”. The only detector is production |
The two error types have wildly asymmetric visibility, and the feedback you receive is almost entirely from the harmless one. That has a nasty second-order effect: the pressure on the classifier will be one-directional. Every week someone will complain that a trivial change was sent for human review. Nobody will complain about the change that was waved through, because the complaint would have to come from a person who never saw it. A classifier tuned by the feedback it actually receives drifts toward permissiveness, and the drift is invisible by the same mechanism that made the original error invisible.
This is not an argument against classifiers. It is an argument that a classifier without a deliberate counterweight will drift, and that the counterweight has to be manufactured, because the environment will not supply it.
Five conditions. The first three are cheap; the fourth is the one nobody does; the fifth is the one that actually protects you.
| Condition | Why | |
|---|---|---|
| 1 | The risk taxonomy is written down and owned by a person | If “risky” is only in the model, you cannot review the policy, only its outputs. Duckbill’s version is a sentence: public API, auth, design system, non-additive schema changes, agent skills |
| 2 | It is opt-in per area, not global | It bounds blast radius to areas someone chose, and makes expansion a decision with a date on it |
| 3 | Irreversibility is in the taxonomy, not just impact | A change that is easy to revert and one that migrates data can have identical blast radius and completely different consequences for being wrong. Classify on how hard it is to undo, which is a property you can determine statically |
| 4 | A sampled audit of the auto-approved path | The counterweight. Pull a random sample of auto-approved changes each week and have a human review them after merge, recording how many should have been escalated. This is the only thing in the list that manufactures the missing signal, and I have not seen an organisation describe doing it |
| 5 | A blameless path for “this should not have been auto-approved” | Because the only other detector is an incident, and incidents make people defensive about the mechanism that caused them |
Note that condition 4 is a specific instance of advice OpenAI gives generally in its own guide for choosing an AI review tool: “Curate examples of gold-standard PRs… Save this as an evaluation set to measure different tools,” and “Define how your team will measure whether reviews are high quality.” The gap is that an evaluation set measures the reviewer; nothing in the published guidance measures the router.
A pipeline is a sequence with an end. A factory has feedback: the output of the process re-enters it as input, without a person carrying it back. That is the whole distinction, and it is why the last two items on the diagram matter more than their size suggests.
Everything in stages 1–8 could be described as a very good pipeline. Stages 8 and 9 turn it into something else: production observations become new work, automatically, and the work re-enters at implementation. The system generates its own inputs.
| Perf Factory | Sevbot | |
|---|---|---|
| Trigger | Continuous — alerts and dashboards | An incident is detected; the bot “wakes up” |
| What it does | “sift through alerts and dashboards, de-duplicate signals, identify real latency regressions, root-cause them and propose fixes” | “Collects context about the incident”; “Determines possible mitigations (but never executes any)”; “Answers devs’ questions (it’s part of the Slack channel)” |
| Human’s role | The fixes are proposed, so they re-enter the normal pipeline and are therefore subject to the stage 6 gate | “An engineer can tell it to apply a specific mitigation” |
| Built on | Agents | “unsurprisingly, it’s also built on top of Codex” |
| Stated future | — | “OpenAI’s goal is to get to the point where Sevbot can take autonomous action when mitigating some outages. The dream is that no humans be woken up outside of their working hours… But as of now, oncall duty is not a thing of the past at the company.” |
The parenthetical is doing enormous work. Sevbot is a proposal engine with a human trigger, and autonomous mitigation is explicitly a goal rather than a state. Any summary of this system that says “an agent handles incidents” has dropped the only clause that determines whether it is a reasonable design.
Notice also the structural consistency with Module 2’s figure: the two places a human retains unambiguous authority are defining intent and taking irreversible production action. Sevbot is the second one. That is not a coincidence — it is the same principle applied at the other end of the pipeline, and it is the strongest evidence that the design has a coherent theory behind it rather than being automation applied wherever it fit.
OpenAI publishes a developer showcase called “SRE Agent
for Incident Response” whose sample app is literally named sev bot. It investigates
incidents in Slack using GitHub and AWS, references past incidents, and proposes remediation. It is
a runnable sample application built on the Agents API, not OpenAI’s internal
system. Its own documentation says “Only a responder can approve a production change,” that
the agent calls propose_rollback but cannot execute it, that “approval records a
decision, not an executed rollback”, that “this app did not execute a rollback”, and
that the first run uses bundled sample telemetry.
Two reasons to care. First, if you go looking for evidence about Sevbot you will find this, and it is easy to mistake a demo for a description of production. Second — and more usefully — the sample encodes the same constraint as the internal system: propose, never execute. When a vendor’s reference implementation and its internal practice agree on a boundary, that boundary is a considered position rather than a limitation of the current model.
| Buys | Costs |
|---|---|
| De-duplication is the real product. “Sift through alerts… de-duplicate signals… identify real latency regressions”. Most alert estates are dominated by duplicates and noise; an agent that collapses them is doing the job a human oncall spends their first twenty minutes on | A new router with the same silent-failure property as the classifier. A de-duplicator that wrongly merges two distinct regressions, or filters a real one as noise, produces no signal. Module 5’s whole argument applies here unchanged, and nobody describes measuring it |
| Work generation without a human courier. A latency regression becomes a proposed fix that re-enters at implementation. The organisation stops losing findings to “someone should file a ticket” | The pipeline can now feed itself faster than it is inspected. Auto-generated work entering a pipeline with an auto-approving gate is the one composition in this system with no human in it anywhere, end to end — if the fix lands in an opted-in area |
| Incident context assembly, which is nearly pure gain. Collecting logs, commits and infrastructure changes is the tedious part of triage and is low-risk to automate because a human still decides | Dependence concentrates. Separately in the same article: OpenAI is so dependent on Codex that in even a minor outage, colleagues’ messages reach the Codex team “at the same time as — or before — automated alerts.” An incident agent built on Codex has a correlated failure with the thing it exists to help debug |
That last cell is the one I would raise in a design review. Building your incident-response agent on the same substrate as your development agent is efficient and creates a dependency that is exactly wrong in the tail: the scenario where you most need incident tooling is the scenario where the shared substrate may be the thing that is broken.
Do first the things whose absence would invalidate the work that follows. In practice this means the unglamorous items come first, because the glamorous ones depend on them and not the reverse. Every step below is a precondition for the ones under it, and the pipeline in Module 2 is the last row, not the first.
| # | Prerequisite | Why here |
|---|---|---|
| 0 | A verification layer you would bet on | Everything in the factory is a loop that terminates when a check passes. If your tests do not fail when the code is wrong, every agent in the pipeline is confidently producing garbage at speed and the routers are routing noise. OpenAI’s own guide puts it the same way: “defining high quality tests is often the first step to allowing an agent to build a feature.” This is not step one because it is virtuous; it is step one because nothing downstream means anything without it |
| 1 | CI throughput sized for ten times the current PR rate | Module 4. A capital project, on your roadmap, that improves nothing visible on the day it lands. Skipping it does not make the pipeline slower; it makes the pipeline impossible, because the agent’s inner loop is fix-rerun-fix |
| 2 | Documentation colocated with code | Module 3. Cheap, independently valuable, and the thing that makes every later stage’s context better. Also the one item here you would do even if you abandoned the whole plan |
| 3 | A written risk taxonomy with a named owner | Module 5. Note this comes before any classifier: you cannot evaluate a classifier against a policy that does not exist, and writing the policy down is most of the value. Duckbill’s fits in a sentence and was enforced with a shell script adding a GitHub label — no model involved |
| 4 | An evaluation set for review quality, before adopting a reviewer | Straight from OpenAI’s guide: “Curate examples of gold-standard PRs… Save this as an evaluation set to measure different tools”, and “Define how your team will measure whether reviews are high quality.” Adopting first and evaluating later means you never find out, because by then the baseline is gone |
| 5 | Post-merge sampled audit of anything auto-approved | Module 5, condition 4. Build this with the auto-approval, never after — if it arrives later you have an unmeasured period you can never reconstruct, and the drift will already have happened |
| 6 | Feature flags and a rollback you have actually used | Agentic deploy supervises a rollout; it does not create the ability to roll one back. If your current recovery story is “revert and redeploy, about forty minutes,” then per-change agent supervision is watching a fire it cannot put out |
| 7 | Alert hygiene sufficient for de-duplication to be meaningful | Module 6. An agent that de-duplicates a noisy alert estate makes the noise tolerable, which removes the pressure to fix it. Doing this before the estate is sane buys you a permanent, automated excuse |
| 8 | Only now: the pipeline | And even then, incrementally, by area, in the order in Module 2 — context assembly first, auto-approval last |
Duckbill, a fifteen-person company, via its cofounder and CEO Mike Julian. Their trigger was recognisably normal: “we found ourselves with 60 open PRs for a team of five… we all had the sudden realization we were looking at two days of just code review.”
What they built before removing review, in their own description:
And the result: PRs merged 353 → 684 (80/wk → 154/wk, +94%); merged within an hour 28% → 45%; human-reviewed median merge time 26 hours, versus 1 hour without.
Three reasons. It is a size you can reason about. It has an actual before and after. And its risk gate is a shell script and a label — no classifier, no model, no agent. The taxonomy did the work; the mechanism enforcing it was deliberately boring.
That is the most transferable finding in this entire course. The valuable half of the risk gate is the taxonomy, and the taxonomy is free. You can have most of the benefit described in the diagram — humans spending their review attention where blast radius is real — with a rule file and a linter, and you can have it this quarter, with none of the silent-failure exposure of a learned classifier. Start there. Move to a classifier when the deterministic rule genuinely cannot express the distinction you need, and when you have the audit from step 5 running.
| Stage | Verdict | Reason |
|---|---|---|
| Docs colocated with code | COPY NOW | Valuable with or without agents; improves every later stage |
| Agent babysits PR to green | COPY NOW | Bounded, reversible, obviously checkable — CI either passes or it does not. Gated only by step 1 |
| AI review as an added lane | COPY NOW | Additive. It comments; humans still decide. Do step 4 first so you can tell whether it is any good |
| Written risk taxonomy + deterministic enforcement | COPY NOW | The Duckbill finding. Free, and it is the valuable half |
| Incident context assembly (propose, never execute) | COPY NOW | Tedious work, human keeps authority, and the boundary is one both OpenAI’s internal and reference implementations hold |
| Specialist review agents | LATER | A context-engineering exercise whose value is unproven even in the source. Real cost, plausible benefit |
| Agentic deploy with self-authored dashboards | LATER | Requires step 6 and a mature observability stack. The agent choosing its own success signals is a bigger delegation than it reads as |
| Learned risk classifier auto-approving merges | LAST | Silent failure mode, no accuracy evidence in any source, and both frontier labs disagree about whether to do it. Last, opt-in per area, with the audit running |
| Autonomous incident mitigation | NOT YET | Nobody has it. It is a stated goal at the company that would most plausibly have built it already |
Nine situations. For each, say who holds decision authority in the pipeline as it actually runs — not as you would design it.
Ten things go wrong. For each, name the detector: does an automated check catch it, does a human notice, or does nothing happen at all until production tells you? This is the tool worth doing twice. The pattern in the answers is the whole argument of the course.
Eight candidate projects. For each: now (cheap, additive, and it unlocks later work), later (real value, but it depends on something you do not have yet), or not yet (nobody can yet show it working).
Start with the input that changes everything downstream: every engineer, researcher, and finance and marketing colleague at OpenAI works with an unlimited token budget. No usage cap on LLMs, the same as at Anthropic. Parallel specialist reviewers on every change, a supervising agent per deployment, nightly passes over the whole codebase — all of that is rational when inference is free at the margin and none of it is obviously rational when it is not.
And the tooling is not tooling you can buy. OpenAI’s internal Codex is substantially more advanced than the product, because it is plugged into pretty much every internal system. Add that they train the model, see its failure modes before anyone else, and can change the model in response to what the pipeline needs — and that the engineers are a heavily selected population. Several of these inputs cannot be purchased at any price.
OpenAI sells Codex. A story about OpenAI running on Codex is the best available advertisement for Codex, and a recruiting pitch to exactly the engineers they want. The people describing the pipeline are leaders whose standing is tied to it working. No sceptic speaks; nobody describes a change that went wrong; there is no failure story anywhere in nine stages. And OpenAI’s engineering guide closes with “reach out to OpenAI. We’re here to help you turn coding agents into real leverage” — explicitly a sales document.
By construction, low-risk changes are the ones where being wrong is cheap. They are also usually the small ones. So the auto-approve gate buys speed on the cheap half of the work while carrying the entire tail risk of misclassification — because every catastrophic outcome of this design lives in the case where something believed low-risk was not. The expected value is a modest, diffuse throughput gain against a rare, concentrated, and structurally undetectable loss. That is a bad shape for a bet, and it is the shape you get whenever you optimise the common case by removing the check that catches the uncommon one.
An organisation can study attention routing, failure asymmetries and readiness ordering for two quarters and ship none of it, while the teams actually getting value run the crude version: turn on an AI reviewer, read the comments for a fortnight, keep it or kill it. Careful analysis is more comfortable than that, and from the outside it is indistinguishable from delay — especially to a leadership team that would rather not commit.
The honest case: every part of this pipeline is a thing that was impossible two years ago and is now routine at one company, and the pattern of such things is that they diffuse. The prerequisites in Module 7 are real today and are exactly the kind of thing that gets commoditised — CI capacity is a purchasing decision the moment someone sells agent-scale CI, and they will. The specialist-reviewer mechanism is plausible on its face and cheap to test. The risk taxonomy is free, as Module 7 itself argues. And the silent-failure argument, while structurally sound, proves too much: human code review has exactly the same asymmetry — a reviewer who skims and approves also produces no signal — and we have run the industry on it for forty years without demanding a sampled audit. Holding agents to a standard we never applied to ourselves is not rigour, it is status-quo bias with a ledger.
| Argument | Why it fails |
|---|---|
| “OpenAI has removed humans from software delivery.” | The article says agentic deploy begins after a human approves the change, that Sevbot never executes a mitigation, and that oncall duty is not a thing of the past. OpenAI’s own guide says engineers own the final review and merge. Module 5. |
| “This is just a normal CI pipeline with AI buzzwords on it.” | The two feedback loops are a category difference: production observations become new work without a human courier, and the work re-enters at implementation. Module 6. |
| “We can’t do this because our models aren’t as good as theirs.” | Model access is a purchase. What stops most organisations is CI throughput, a verification layer worth betting on, and a written risk taxonomy — and the measured example in this material is a fifteen-person company whose risk gate was a shell script. Modules 5 and 9. |
Write down what “risky” means in your codebase, in one paragraph, and have one person own it. Enforce it with the dumbest mechanism that works. Then look at what your engineers currently spend review attention on and ask whether that paragraph is where it is going. You will have implemented the valuable half of the most contested box in the diagram, this week, with no agent involved — and you will have a policy you can actually evaluate a classifier against, on the day you decide you want one.