A developer at a twelve-year-old codebase emailed the creator of Claude Code asking him to settle an argument. Boris Cherny answered, then published both halves. This is a close reading of what each actually said, what the answer costs to follow, the one question in the email it does not address — and the strongest arguments against it.
Primary source, read in full: Boris Cherny’s post, 10 Sep 2026 (both email screenshots and his reply) · corroborated by Business Insider, Henry Chandonnet, 12 Sep 2026
Anyone who has to hold a position on AI-written code in front of a team that is already split about it. Two camps have formed, and both have a real argument: one says AI-generated code is fine if it is reviewable and the submitter can explain it, the other says the review cost has quietly overtaken the generation saving. This course takes the exchange that put the question to the person who built the tool, reads both halves closely, prices the answer he gave, and sets out what the answer leaves alone.
Quotations are short, exact and attributed — a few sentences from each side. No view is attributed to anyone who did not state it, and the correspondent is treated as anonymous because Cherny redacted his name, employer and signature before publishing.
Short on time? Module 1 for what was actually published and by whom, and Module 6 for the question the reply leaves alone. Those two carry the course.
On 10 September 2026 at 6:10 PM, Boris Cherny — creator and head of Claude Code at Anthropic — posted two screenshots of an email he had received, with the note:
“Every day, I get a lot of of emails and messages like this one. I try to respond to as many as I can. Sharing my response below, for anyone else in a similar situation. What do you think?” Boris Cherny (@bcherny), 10 September 2026 — typo in the original
The next day he posted his reply to the email as a follow-up in the same thread. The parent post has since passed 750,000 views; the reply, 362,000. The email’s subject line, visible in the screenshot, is “What to do about slop?”, and the thread header shows 4 messages — so what was published is one exchange out of a longer back-and-forth.
This was not a private email leaked by its sender. The recipient published it. Cherny posted an unsolicited email he received, together with his own reply. Business Insider’s account agrees: “On Thursday, Anthropic’s Boris Cherny posted an email a developer wrote to him on X.”
That matters for how you read the whole thing. Nobody was exposed by an adversary. The asymmetry that remains is a different one: the person who published controls the framing, chose which of the four messages to show, and is the one the excerpt makes look good.
Cherny redacted, in the images themselves, every detail that would identify the sender:
| Redacted | Where |
|---|---|
| The sender’s name | The mail app’s header bar, and the salutation of the reply — which reads “Hey ████,” |
| The sender’s employer | Two blacked-out spans in the opening sentence (“I’m a dev at ████”) |
| The sender’s signature block | Blacked out above “Sent from my iPhone” |
The redactions are the most creditable part of how this was handled, and they are why the developer is treated here as an anonymous correspondent. He did not choose to be published. He was published carefully, but he was still published, and nothing in the post says whether he was asked first. The phrase “for anyone else in a similar situation” points to a public-FAQ motive rather than exposure — though that is a reading of intent, not something either party stated.
It will not try to work out who the developer is. Cherny went to the trouble of blacking out his name, his employer and his signature; reversing that would be a strange thing to do in a course about professional judgement. Everything below treats him as “the developer” and draws only on what he himself wrote in the published screenshots.
He describes himself in one line: a developer at a company he has been with for twelve years, where, in his words, “Now I’m just trying to keep it running.” He says the shift to agentic development “is causing some friction” and that there are two views in the organisation. He is not asking whether AI coding is good. He is asking someone with standing to settle an argument his team is already having.
That self-description does a lot of work and is worth holding on to for the rest of the course. This is maintenance of a mature system by someone who has been there long enough to own the consequences — not greenfield, not a startup, not a team with spare capacity.
“We create code similar to before, accelerated by AI. We may not review all of it, but it should be reviewable. The person submitting the code should be able to explain it. It should be built in such a way to make it as easy to maintain – or maybe easier – than before, by humans or agents.”
The developer’s words, from the published screenshot
“Vibe code: We treat the code as a black box. We don’t worry about it. Just check the output.”
The developer’s words, from the published screenshot
Notice how carefully View 1 is drawn. It is not “review everything” — he concedes up front that “we may not review all of it.” The bar he proposes is reviewability and explainability by the submitter, not review. That is a more sophisticated position than the caricature of the cautious side, and it survives most of the standard objections to “just review the code.”
Strip the View 1 / View 2 framing away and the email is one economic claim: generation got cheap and verification did not. Every other complaint follows. People produce more than can be checked; the check is the expensive step; so the organisation drifts toward View 2 not because anyone argued for it but because it is the only thing that scales.
The developer buried his strongest point in the fourth paragraph and led with a taxonomy. That is a very normal thing to do, and it shapes the answer he got. Module 6 is about what happened to claim A.
The word appears in the email’s subject line and nowhere else in either message — neither the developer nor Cherny defines it. What follows is this course’s working definition, offered to make the rest legible rather than attributed to either of them.
The everyday use of “slop” means low-quality generated output. That reading makes the email incoherent — if the code were simply bad, there would be no disagreement to broker, and View 2 would have no advocates. The version that fits is narrower:
Slop is output produced faster than the receiving system can verify it.
Three properties follow. It is a rate problem, not a quality problem — the same change is slop or not depending on what capacity exists to check it. It is relational — defined against the team, not the artefact. And it is invisible at the point of production, because the person generating it is the one least positioned to notice the backlog forming behind them.
This is why “the code is fine” and “this is slop” are not contradictory, and why both sides of the developer’s argument can be reporting honestly. The View 2 advocate says the output passes. The View 1 advocate says nobody can now explain why it passes. Both are true.
Writing code was expensive; reading it was cheap by comparison. Review scaled because production was slow. A reviewer could keep up because the bottleneck sat upstream of them.
Quality control could therefore be a gate: everything passes through, because the flow rate was survivable.
Production collapsed in cost; comprehension did not. The reviewer is now the slowest step in the pipeline, and the gate becomes a queue.
A gate with a queue behind it is not a gate. It becomes a formality, and the organisation quietly reclassifies as View 2 without ever deciding to.
That drift — toward View 2 by congestion rather than by argument — is the most plausible reading of what the developer means by “friction.” He never puts it in those terms; this is the course’s reconstruction.
He opens: “I think there is room for both.” The developer asked him to pick a side and settle it; he declines the frame and replaces it with a split.
“Prototypes and other throw-away code can be treated as totally black box. If you’re going to throw it away anyway, and if the blast radius of it breaking is low, it doesn’t need to be perfect.” Boris Cherny, 11 September 2026
This is the most portable idea in the whole exchange and it costs nothing to adopt. The criterion is consequence of failure × expected lifetime, and both are properties of the change rather than of who or what produced it. It gives the developer something his two-view framing did not have: a way for both camps to be right about different code on the same day. Module 8 makes you apply it.
“Production code written by Claude should have a higher bar than if it was written by a human.” Boris Cherny, 11 September 2026
The head of the product is saying his product’s output should be held to a stricter standard than a human’s. He was under no commercial obligation to say that, and it is the opposite of the claim a vendor is usually accused of making. Any fair reading of this exchange has to credit it.
It also quietly concedes the developer’s premise. You do not need a higher bar for output you trust equally.
He lists what Anthropic runs to hold that bar: “lots of lint rules, lots of tests, Claude-driven end to end tests, Claude-powered fuzzers running daily, automated code reviews and security reviews, automated code refactoring, and so on.” And the consequence of not having them: “Without these, you can end up with a mess that is hard to maintain down the line.” Module 5 is about what that list costs.
“Your job is to hold the bar on code
quality.” If Claude’s output misses it, he offers, in order: use the latest frontier
model (naming Opus 5 or Fable 5.1); raise effort to high or xhigh; invest in
CLAUDE.md and skills to teach Claude the codebase; and failing those, steer more, have
Claude pay down accumulated debt and rewrite the codebase, “Or, wait for the next
model.”
Move 2 and Move 3 are principles: they hold whatever tooling you use, they are checkable, and you could adopt them on a competitor’s model tomorrow. Move 5 is largely product advice — three of its four suggestions route through buying more of, or configuring more of, the vendor’s own offering.
That is not a gotcha. He was asked a question by a user of his product and answered it from where he sits. But the two piles have different shelf lives, and a reader deciding what to take from this should know which pile each item is in.
Cherny’s answer to “how do I stop this becoming a mess” is a list of seven things Anthropic runs. Every one is a real control and the list is honest — it is what a well-resourced platform team builds. The gap is between that list and the person who asked, who described himself as maintaining a twelve-year-old system and “just trying to keep it running.”
| Guardrail (his words) | What it costs, realistically | Reach of a mature codebase |
|---|---|---|
| “lots of lint rules” | Cheap to add, expensive to adopt: on a large old codebase the first run produces thousands of violations, so you need baselining and a ratchet | highest value per hour of the seven |
| “lots of tests” | The classic cost. On legacy code the blocker is testability, not test-writing — you often cannot test without refactoring first | the bit agents genuinely help with |
| “Claude-driven end to end tests” | Needs a working E2E harness and an environment to run it in. That infrastructure is the cost; the authoring is not | presumes a test environment many maintenance teams lack |
| “Claude-powered fuzzers running daily” | Continuous compute plus someone to triage findings. Fuzzers generate work; unowned, they become a second backlog | needs a standing owner, not just a budget |
| “automated code reviews and security reviews” | Cheapest to switch on of the seven. Per-PR cost, no infrastructure | the other high-value item |
| “automated code refactoring” | Cheap to run, expensive to trust — and trusting it is exactly what is in dispute | see the box below |
Four of the seven controls are themselves AI-driven: Claude writes the E2E tests, Claude fuzzes, Claude reviews, Claude refactors. The developer’s worry is that he cannot verify AI output at the rate it arrives. A remedy substantially composed of more AI output does not dissolve that worry — it relocates it, from code he cannot check to controls he cannot check.
This is not fatal. Verification is genuinely easier than generation in many cases: a failing test is a discrete, checkable signal in a way that a thousand-line diff is not. But the reduction has to be argued, and the reply asserts rather than argues it. Notice too that the two items I marked “within reach” — lint and automated review — are the two least dependent on trusting a model’s judgement.
Lint rules with a ratchet, and automated review on every change. Both are deterministic or near-deterministic, both are per-change rather than per-environment, neither needs a test harness you do not have, and neither creates a backlog that needs an owner. They also bite hardest on exactly the failure mode in question: volume arriving without anyone reading it.
| The developer’s claim | What the reply does with it |
|---|---|
| The two-view disagreement — which side is right? | “There is room for both,” split by blast radius. A clean, usable answer to the question as posed. |
| (A) Verification now costs more than generation | The reply prescribes more verification — tests, fuzzers, reviews — without engaging the claim that verification is the expensive side of the ledger. |
| (B) The layoff dynamic pushing toward volume | Answered technically: “Your job is to hold the bar.” The developer’s point was that holding the bar is a political position inside his org, not only a technical one. |
| (C) Agents miss the simpler design | “Use the latest frontier model” and “increase effort” are plausible responses to an agent not finding a better design. Whether a stronger model proposes relaxing the spec — his actual example — is untested. |
The reply’s prescription is more verification. The email’s thesis is that verification is the constraint. If the thesis holds, the prescription lands on the resource already exhausted — which is the definition of advice that is correct and unusable.
There is a real answer available in the guardrail list, and I think it is the one Cherny means: automation moves verification from human-time to machine-time, so the constraint moves too. A daily fuzzer does not consume the reviewer’s attention until it finds something. That is a genuine dissolution of the asymmetry, not a dodge — but it is never stated, and the developer, reading the reply, would have to reconstruct it himself.
That reconstruction is mine, not his. He may have meant something else, or simply been answering the question he was asked in the space an email allows.
He was asked one question — “shed some light on how to handle the disagreement” — and he answered exactly that, for free, to a stranger, within a day. Faulting a short email for not also solving the economics of software verification and his correspondent’s internal politics would be unreasonable. The point of this module is not that the reply is deficient. It is that a reader who takes it as a complete answer to the email will have taken one third of one.
This exchange is two practitioners reasoning from experience. That is legitimate, and neither cites evidence. But the developer made a falsifiable economic claim, and there is measurement that bears on it — so a course about the exchange should put it next to them. None of what follows is attributable to either party.
| Finding | What it supports | What it does not |
|---|---|---|
| METR’s randomised trial: experienced open-source developers were 19% slower with AI assistance, while believing they had been 20% faster | The developer’s asymmetry, in its sharpest form: the cost lands on the person doing the work and is invisible to them. A team can drift to View 2 while sincerely reporting speedups. | It is experienced devs on mature repos they know well — close to the email’s situation, but a narrow population. It says nothing about prototypes, where Cherny’s black-box concession applies. |
| DORA 2025: throughput up and stability down in the same population | Exactly the shape of the complaint: more shipped, less reliably. Supports “the gate became a queue.” | Correlational across a survey panel. It does not isolate AI as the cause, nor identify which practices separate the teams that avoided it. |
| GitClear’s 623M-change corpus: refactoring down 70%, duplication up 81% | The best direct evidence for maintainability decay at volume — and for the developer’s (C), that the simpler design is not being found. | Repository metrics, not outcomes. Duplication is a proxy for maintenance cost, not a measure of it, and the corpus cannot attribute changes to a cause. |
It strengthens the developer: his claim is not a temperamental preference for caution, it is the reported pattern at population scale, and the METR result supplies the mechanism for why the disagreement persists — the people generating volume are not lying about feeling faster.
It also does not refute Cherny, and this is the part people skip. None of the three studies measures teams running his guardrail stack. They measure the counterfactual he is arguing against — AI-accelerated development without lint ratchets, automated review, daily fuzzers. His claim is conditional: with these controls you avoid the mess. The evidence describes the world without them.
So the honest summary: the disease is well documented; the prescription is untested at population scale. That is not a criticism of the prescription. It is the state of the knowledge.
Cherny’s rule has two inputs: is this going to be thrown away, and what is the blast radius if it breaks. Ten changes below. For each, decide whether black box is acceptable or the higher bar applies. Some are deliberately underdetermined — the rule is good, not total, and finding its edges is the point.
Everything under an heading is an argument against a position, not a view held by the person who holds that position. Where the source contains material that answers an objection, it is quoted. Where I construct a reply neither of them made, it is marked and is mine.
The correspondent described himself as maintaining a twelve-year-old system and “just trying to keep it running.” The remedy is seven controls including daily fuzzers, Claude-driven E2E tests and automated refactoring — a platform-team programme. Advice that is correct for Anthropic and unreachable for the person asking has not answered the person asking.
The developer’s problem is that the bar cannot be held at the volume arriving. Telling him it is his job to hold it settles ownership, which was not in dispute, and leaves the constraint untouched. It also lands awkwardly next to his point about colleagues generating volume he then has to absorb — on that reading the sentence assigns him the cost of other people’s throughput.
It is the last item, and it is unfalsifiable in the moment: if quality is inadequate today, deferring to an unreleased model is a prediction the reader cannot evaluate, act on, or plan around. For someone maintaining a production system now, it functions as a deferral rather than a step.
Four of the seven guardrails are model-driven. Where the authoring model and the reviewing model come from the same family, errors they are jointly disposed to make survive the check. The control and the thing controlled are correlated, and the developer’s original worry — that he cannot verify what arrives — is relocated rather than resolved.
Both inputs are assessed when the code is written. Neither is re-checked. The migration script that becomes a scheduled job, the demo that becomes the onboarding path, the internal tool that acquires external users — each was correctly classified as throwaway at the time, and no part of the rule ever prompts a re-evaluation. The failure is invisible because the original decision was right.
“The person submitting the code should be able to explain it” is a lower bar than full review, but it is still a per-change human cost that rises linearly with volume. If generation is 10×, explainability-on-demand is 10× the explaining. View 1 is more defensible than the caricature, and it may still be a slower version of the same wall.
Verification costing more than generation may be a property of this moment — immature tooling, teams without harnesses, controls not yet built — rather than a standing law. If so, building the controls is exactly right and the discomfort is the cost of not having built them yet.
“Many people, who may be at risk of being laid off, who can’t do 1. But they can generate a lot of 2” attributes a technical position to job insecurity. Even if the correlation is real, it makes it harder for a View 2 advocate to argue on the merits without appearing to defend their own position — which is corrosive to exactly the disagreement he asked to have resolved.
| The objection | Why it fails |
|---|---|
| “He sells the tool, so the answer is marketing.” | Genetic fallacy, and the text runs the other way: he argues his product’s output should meet a higher bar than a human’s, and warns you can “end up with a mess that is hard to maintain.” The disciplined version is objection 3, which is about one specific line. |
| “It shows AI-generated code is simply bad.” | Neither party says this. View 2 has advocates precisely because the output passes; the developer’s complaint is about verifiability at volume, not defect rates. Attacking quality misses the argument entirely. |
| “A private email should never have been published.” | Half-right at most. The recipient published it, having blacked out the sender’s name, employer and signature. The live question is consent, not exposure — and nothing in the post says whether he asked. |
The failure mode with a short, widely-shared exchange is drift: a paraphrase hardens into a quote, and analysis gets attributed to the participants. Twelve statements. For each, decide whether it is the developer’s, Cherny’s, this course’s own analysis, or something nobody in this material said.