The Context Budget

Skills, CLAUDE.md, memory, retrieval, MCP and subagents are not six features. They are six competing answers to one question — what should the model see right now — and they differ on when the decision gets made, who makes it, and what it costs every turn. Written for the person who has to walk into a room of engineers and leaders and be credible about all of it, then go and actually get it adopted.

12 modules
3 interactive tools
37 quiz questions
~70 min
Prices checked 14 Sep 2026

Primary sources: models overview · prompt caching · Agent Skills · API and data retention · issue #14882 · managed MCP · Enterprise-Managed Authorization · MCP authorization security · Claude Code analytics · champion kit

Who this is for, and what it refuses to do

This is for someone arriving to help an engineering organisation adopt AI — the person who will be asked, in the same week, what this costs, who owns it, whether it is compliant, and why the pilot has gone quiet. Two things follow. The mechanics are not optional: the token arithmetic and the mechanism taxonomy in Modules 1–8 are what separate someone credible in a room of engineers from someone hand-waving, and hand-waving is detected in about ninety seconds. And the mechanics are not the job: Modules 9–11 are about sequencing, what to do first, how to tell real adoption from the appearance of it, and which of the arguments you will meet are actually correct.

You should finish able to lead a conversation about this rather than survive one — which means knowing which objections to concede in the first thirty seconds so that the rest of what you say lands. It does not explain what a token is beyond what the arithmetic needs.

Prices are as of 14 September 2026. They move, and the arithmetic here is built so you can redo it with whatever the numbers are when you read this.

Short on time? Module 2 for where the money actually goes, Module 8 for the compliance exclusions that are mutually exclusive, and Module 10 for sequencing and what the adoption metrics cannot tell you. Those three change decisions. Module 11 is where you find out whether any of it stuck.

Course Modules

  1. One question, three axesThe spine
  2. Where the money actually goesEconomics
  3. Why a bigger window doesn’t helpScarcity
  4. Memory, context and retrieval are not the samePrecision
  5. The mechanisms, pricedSelection
  6. MCP gatewaysControl
  7. When the documented cost isn’t the real costMeasured
  8. Ownership, review and mutual exclusivityEnterprise
  9. Practices that earn their placePractice
  10. Driving adoptionThe job
  11. Three decisions, scoredInteractive
  12. CounterargumentsBoth sides
1

One question, three axes

The frame that turns a glossary into a decision procedure
By the end of this module you will
  • Be able to place any context mechanism on three axes without memorising a list
  • Know the failure mode that explains most bad mechanism choices

The question

A language model is a function from a context window to a completion. It has no state between calls. Everything an enterprise buys, builds or governs around it — skills, CLAUDE.md files, memory products, retrieval pipelines, MCP servers, subagents, fine-tunes — exists to answer one question:

The question

What should be in the context window at this moment, and what did it cost to put it there?

Treating these as six features to evaluate separately is the single most expensive framing error in enterprise AI adoption, because it produces six procurement conversations, six owners and six budgets for what is one resource-allocation problem.

The three axes

1 · DECISION TIME
When is it decided that this content belongs in context? Fine-tuning decides at training time. A CLAUDE.md decides at authoring time and is applied at load time. A skill decides at invocation time. Retrieval and MCP decide at turn time. A subagent decides at call time and then throws the decision away.
2 · DECIDER
Who holds the authority? The organisation (managed settings, an org-tier plugin), a team (a repo’s CLAUDE.md), an individual (a personal skill), or the model itself (it chooses which skill to trigger, which document to retrieve, which subagent to spawn). That last case is the one governance frameworks routinely fail to account for.
3 · PER-TURN COST
What does it cost every single turn, whether or not it is used? This is the axis tokens price, and the one that separates mechanisms that look equivalent. A CLAUDE.md is paid on every turn of every session forever. A skill body is paid only in sessions that trigger it. A subagent’s context is paid once and discarded.
The failure mode this predicts

Most bad mechanism choices are decision-time mismatches: the mechanism’s decision moment does not match how often the answer actually changes.

  • Putting volatile facts in a CLAUDE.md — deciding at authoring time something that changes weekly — produces a file nobody trusts and everybody pays for on every turn.
  • Fine-tuning on a policy that changes quarterly — deciding at training time something that changes faster than you can retrain.
  • Retrieving, per turn, a coding standard that has not changed in a year — paying turn-time cost for an authoring-time fact.

The rule that falls out: match the decision moment to the rate of change, and match the decider to whoever is accountable when it is wrong. Module 5 applies it mechanism by mechanism; Module 11 makes you do it.

Takeaways
  • One question: what should be in context now, and what did it cost to put there
  • Three axes: decision time, decider, per-turn cost
  • Most bad choices are decision-time mismatches — mechanism moment vs rate of change
  • In several mechanisms the decider is the model, which most governance frameworks do not model
2

Where the money actually goes

Real prices, real arithmetic, and a result most budget models get wrong
By the end of this module you will
  • Know the current prices and the three multipliers that change them
  • Be able to model an agentic loop and see where the cost concentrates
  • Know why the intuition “input is cheap, so context is cheap” is wrong in a loop

The numbers, checked 14 September 2026

ModelInput /MTokOutput /MTokCache readContextMax output
Claude Fable 5.1$10$50$0.25 (2.5%)1M128K
Claude Opus 5$5$25$0.50 (10%)1M128K
Claude Sonnet 5$2$10$0.20 (10%)1M128K
Claude Haiku 4.5$1$5$0.10 (10%)200K64K

From the models overview and pricing pages. Three multipliers stack on top: cache writes cost 1.25× base input at the default 5-minute TTL and 2× at 1-hour TTL; cache reads cost 10% of base input (2.5% on Fable 5.1 and Mythos 5.1); and the Batch API is 50% off. US-only inference is 1.1×. Output is priced at 5× input across the lineup — hold on to that ratio, it drives everything below.

The arithmetic that matters

Take a realistic agentic task on Opus 5: a coding agent doing 20 turns, with a 30K-token stable prefix (system prompt, CLAUDE.md, tool definitions) and roughly 2K tokens of new content and 1K tokens of output per turn. — the prices are real, the shape is illustrative.

Stable prefix, 30K tok — no caching, ×20 turns @ $5/MTok$3.00
Accumulated history, ~grows 3K/turn, resent every turn (~570K tok total)$2.85
Output, 1K × 20 = 20K tok @ $25/MTok$0.50
Uncached total$6.35

Now switch on prompt caching for the stable prefix.

Prefix: 1 write @ 1.25× ($6.25/MTok) + 19 reads @ $0.50/MTok$0.47
History, cached incrementally as it stabilises$0.35
Output, unchanged — caching does nothing for output$0.50
Cached total$1.32
The counter-intuitive result

Caching cut the bill by 79%. And look at what is left: output is now 38% of the total, up from 8%. It did not get more expensive — everything around it got cheaper.

This is the finding budget models miss. Teams optimise input because input is what they can see growing, and input is the part that caching already fixes for free. Once caching is on, the remaining lever is output volume — how verbose the model is, how many turns the loop takes, how much it re-reasons. Those are controlled by effort settings, prompt design and loop architecture, not by trimming context.

The two things that actually drive an agentic bill

1. Repeated context (fixable, free)

Every turn resends the whole history. A 20-turn loop with a 30K prefix pays for that prefix 20 times. This looks alarming and is almost entirely solved by prompt caching, which is a configuration change, not an architecture change.

The trap: caching is a prefix match — any byte change anywhere in the prefix invalidates everything after it. A timestamp in the system prompt, a re-ordered JSON key or a varying tool list silently drops your hit rate to zero, with no error. Check usage.cache_read_input_tokens; if it is zero across repeated requests, something is invalidating.

2. Output tokens (expensive, architectural)

Priced at 5× input, never cacheable, and in a loop the output of turn n becomes the input of every turn after it — so a verbose turn is billed once at output rates and then repeatedly at input rates.

That compounding is the real cost of verbosity, and it is why effort settings and loop design matter more to an agentic bill than context trimming. A model that solves the task in 8 turns instead of 20 saves more than any context optimisation available to you.

The practical ordering for an enterprise: turn on caching (free, large), then batch anything latency-tolerant (50%, free), then tune effort per route (measurable), and only then argue about context contents. Most organisations do these in exactly the reverse order because context contents is the part that feels like engineering.

Takeaways
  • Checked 14 Sep 2026: Opus 5 $5/$25, Sonnet 5 $2/$10, Haiku 4.5 $1/$5, Fable 5.1 $10/$50. Output is 5× input everywhere
  • Cache write 1.25× (5-min) or 2× (1-hr); cache read 10% of input (2.5% on Fable 5.1); batch 50% off
  • Caching cut the modelled loop 79% — and left output as 38% of what remained
  • Cheapest-first ordering: caching → batch → effort tuning → context contents
3

Why a bigger window doesn’t help

1M tokens is real, and it does not dissolve the problem
By the end of this module you will
  • Be able to answer “can’t we just put everything in the window now?”
  • Know the three separate constraints a bigger window does and does not relax

Fable 5.1, Opus 5 and Sonnet 5 all carry 1M-token context windows; Haiku 4.5 carries 200K. The docs put 1M at roughly 555,000 words on the current tokenizer. That is a large fraction of an internal knowledge base.

ConstraintDoes a bigger window relax it?
Capacity — will it fit? YES Genuinely solved for most enterprise documents. This is the constraint people mean, and it is the one that is gone.
Cost — what does it cost per turn? NO Strictly worse. Filling 1M tokens on Opus 5 costs $5 per turn uncached, or $0.50 per turn on cache reads. A 20-turn session on a full window is $10–$100 depending on caching. Capacity became affordable to use long before it became affordable to fill.
Attention — will the model actually use it? NO Relevance density still matters; a fact competing with 900K tokens of other material is not in the same position as one competing with 9K. Long-context benchmarks have improved markedly, but "it fits" and "it is attended to" remain different claims.
Latency — how long will it take? NO More input is more time to first token. For an interactive tool this is a product constraint, not an accounting one.
Governance — should this content be there? NO A bigger window makes it easier to include material nobody approved. Capacity removes the accidental forcing function that used to make someone choose.
The line worth taking into a planning meeting

The window got bigger; the budget did not. Context stopped being scarce in the capacity sense and remains scarce in the economic sense — and the economic constraint is the one that scales with your usage rather than sitting fixed. An organisation that treats 1M tokens as “we no longer need to choose” has converted a design constraint into a line item that grows with adoption.

Takeaways
  • 1M tokens on Fable 5.1, Opus 5, Sonnet 5; 200K on Haiku 4.5 — roughly 555k words
  • Bigger windows solve capacity and worsen cost and latency
  • They do not guarantee attention, and they actively weaken governance by removing the forcing function
  • Scarcity moved from capacity to economics — and economics scales with adoption
4

Memory, context and retrieval are not the same

Three different things, bought as if they were one
By the end of this module you will
  • Be able to tell a vendor which of the three they are actually selling
  • Know which one your compliance team needs to care about most
 ContextRetrievalMemory
What it isThe tokens in this one requestA lookup that selects tokens to put in this requestState that persists across requests and is written back
LifetimeOne callThe index lives; each result is per-callIndefinite, until deleted
Who decidesWhoever assembled the requestUsually the model, via its queryUsually the model, via what it chooses to write
Cost shapePer turn, every turnPer hit, plus index upkeepStorage, plus per-turn cost of whatever is loaded back in
Compliance surfaceTransient — governed by retention termsYour index; your data residencyA new persistent store of model-selected content
The distinction that costs money

“Memory” is the word doing the most work in enterprise AI sales, and it is routinely used for all three. Ask one question to disambiguate: “After this conversation ends, what is written down, where does it live, and who chose what went in it?”

  • Nothing is written — it is context. A per-turn cost and no new data surface.
  • Nothing new is written, but a pre-existing index is read — it is retrieval. Your governance problem is the index, which you already have.
  • Something new is written, and the model chose what — it is memory, and you have just acquired a persistent store whose contents no human curated.

Only the third creates a genuinely new compliance object. That is why it matters that the three are priced and procured as if interchangeable.

The retention docs make the practical version of this concrete: Claude Managed Agents — a stateful, memory-shaped product — is called out as not ZDR-covered precisely because “Claude Managed Agents is a stateful resource; session transcripts persist until you delete them.” Statefulness and zero retention are the same tradeoff wearing two names.

Takeaways
  • Context = tokens in one call. Retrieval = a lookup feeding that call. Memory = state written back
  • The disambiguating question: what is written down, where, and who chose it
  • Only memory creates a new persistent store of model-selected content
  • Managed Agents is outside ZDR because it is stateful — the tradeoff is structural, not a gap
5

The mechanisms, priced

Placed on the three axes, with the case for not using each
By the end of this module you will
  • Know each mechanism’s decision time, decider and per-turn cost
  • Know, for each, the situation where reaching for it is over-engineering
MechanismDecides atDeciderPer-turn costOver-engineering when…
CLAUDE.mdAuthoring timeWhoever owns the repoEvery turn, always. Sits in the cached prefix, so ~10% of input rate after the first write — but never zero, and never conditionalThe content is conditional, volatile, or long. Everything in it is a tax on every session including the ones it is irrelevant to
SkillInvocation timeThe model, matching your descriptionDescription always; body only when triggered The procedure is one paragraph, or it always applies (then it is a CLAUDE.md), or it never applies (then delete it)
Retrieval / RAGTurn timeThe model, via its queryPer hit, plus index maintenance you own foreverThe corpus is small enough to sit in the prefix, or static enough to be a skill. A RAG pipeline for 40 pages of policy is a maintenance obligation bought to solve a formatting problem
MCP serverTurn timeThe model, choosing a toolTool definitions every turn; results per callThe data is static. MCP earns its keep on live external state. Tool definitions are permanent prefix cost — a server exposing 40 tools taxes every turn
SubagentCall timeThe model, deciding to delegateIts own full context, paid once and discardedThe sub-task is short. You are paying a fresh context assembly to save a context you were not close to exhausting
Fine-tuningTraining timeWhoever ran the jobZero per turn — the only mechanism with noneAnything that changes. Zero per-turn cost is bought with maximum change latency, and most enterprise knowledge changes faster than a retrain cycle
The structural insight

Read the “decides at” and “per-turn cost” columns together and they are the same column. The earlier a mechanism decides, the cheaper it is per turn and the staler it is allowed to get. Fine-tuning decides earliest and costs nothing per turn; retrieval decides latest and costs on every hit. There is no mechanism that is both current and free, and there never will be, because currency is purchased with per-turn work.

So the selection rule is not “which is best” but: how often does this content change, and what is the cheapest decision moment that still catches the change?

The one that surprises people: CLAUDE.md is not free

Because it is a text file in a repo, a CLAUDE.md feels like documentation, and organisations grow them the way they grow READMEs — additively, with no deletion process. But it is loaded into every session, so every line is billed on every turn of every session by every engineer, forever, whether or not it is relevant.

At ~10% of input rate on cache reads, a 5,000-token CLAUDE.md costs about $0.0025 per turn on Opus 5. Across 200 engineers doing 100 turns a day that is about $50 a day, $12,500 a year, to carry one file. That is not a crisis. It is, though, a recurring cost with no owner and no review cadence, which is exactly the profile of costs that grow unexamined.

Takeaways
  • Decision time and per-turn cost are the same column: earlier decisions are cheaper and staler
  • No mechanism is both current and free — currency costs per-turn work
  • Selection rule: cheapest decision moment that still catches the rate of change
  • A 5K-token CLAUDE.md across 200 engineers models at ~$12.5K/year — small, recurring, and usually unowned
6

MCP gateways

The chokepoint the protocol doesn’t give you, what exists to buy, and what it costs you to install one
By the end of this module you will
  • Be able to say what an MCP gateway does that a direct MCP connection cannot
  • Know which of those five jobs the platform already does natively and which you assemble yourself
  • Be able to name real implementations rather than describing an abstraction
  • Be able to argue the case against a gateway without being talked out of it

Why MCP is the mechanism that needs a control point

Go back to the three axes. MCP is the turn-time mechanism whose decider is the server — the tool list, the tool descriptions and the returned content are all authored by someone outside your organisation and injected into your context at request time. That combination is unique among the six. A CLAUDE.md is yours. A skill is yours. A retrieval index is yours. An MCP server is a third party with a live seat inside every turn.

Anthropic states the position plainly: it “reviews connectors against its listing criteria before adding them to the Anthropic Directory, but doesn’t security-audit or manage any MCP server.” The default is also explicit: “By default, anyone running Claude Code can connect any MCP server they choose.”

The five jobs a gateway is supposed to do

JobWhat goes wrong without it
Credential custodyEvery developer holds a long-lived token for every upstream system on their laptop. Offboarding means finding them. The managed-MCP docs warn directly: “Any user on the machine can read this file, so don’t store API keys or other credentials in env blocks.”
A single audit pointYou can reconstruct what a model said but not what it did to your systems, and each client keeps its own partial record in its own format.
Rate limitingAn agentic loop is a tool-call amplifier. One bad prompt becomes ten thousand calls against a system whose owner never agreed to that load.
Allowlisting which servers existThe set of third parties with a seat in your engineers’ context windows is whatever each engineer chose, and nobody has the list.
Version pinningA remote server changes a tool description overnight and every agent in the company behaves differently the next morning, with no diff, no release note and no rollback.
Version pinning is the one that should worry you most

The other four have recognisable enterprise analogues — they are secrets management, logging, quotas and software allowlisting, and your organisation already has opinions about all four. Version pinning is different, because the artifact that changes is prose the model reads. A tool description is a prompt. Editing it changes model behaviour across every workflow that touches it, and it passes through none of the review that a code change would. You have a dependency you cannot pin, cannot diff and cannot roll back.

What the platform already does, and where it stops

Claude Code ships real administrative control over which MCP servers load. It is worth knowing precisely, because it is more than most people assume and less than a gateway.

ControlWhat it does
managed-mcp.jsonDeploys a fixed set. Exclusive: in cloud sessions on a host where the file is deployed, Claude Code “starts with the managed servers only and skips the claude.ai connectors and other servers the cloud host delivers”
managedMcpServersProvides org servers alongside what users add. Takes precedence over a same-named local, project, user or plugin server. Requires v2.1.259+
allowedMcpServersAllowlist by serverUrl (* wildcards), serverCommand, or serverName (exact match only, no wildcards)
deniedMcpServersDenylist, same entry shapes. “Nothing overrides a denylist match”
allowManagedMcpServersOnlyMakes the allowlist authoritative. Without it, allowlists from every settings scope merge — including the user’s own ~/.claude/settings.json
OTEL_LOG_TOOL_DETAILS=1Records MCP server and tool names in OpenTelemetry tool events, so you can aggregate which servers users actually connect to
Three specifics that decide whether your policy is real

1. A soft allowlist is user-broadenable. Set allowedMcpServers without allowManagedMcpServersOnly: true and every settings scope merges, so a developer can widen your allowlist in their own settings file. You will still see an allowlist in your configuration management and believe it is enforcing something.

2. Unset and empty are opposite. An unset allowedMcpServers means all servers allowed. Set to [] it means no servers allowed. The distance between “we haven’t configured that” and “we configured that to empty” is the whole policy.

3. A client upgrade silently changed enforcement. Before v2.1.259, every managed server had to pass the allowlist. After it, allowedMcpServers no longer applies to managed-mcp.json servers except where ${VAR} expansion is used. The docs say so explicitly: servers you had suppressed that way “start loading on each user’s first launch of v2.1.259 or later, with no prompt or notice.” Whatever else you take from this module, take this: on a fast-moving client, your enforcement posture has a version number.

Where it stops. These controls decide which servers load. They are client-side and per-client. They do not hold credentials on your behalf, do not rate limit, do not sit in the path of a tool call to block it, and — the significant gap — there is no version pinning: a server is identified by name, URL or command, never by a version or a digest. And “Claude Code doesn’t have a built-in MCP server registry that users can browse and install from.”

The distinction to carry into the room: telemetry tells you what happened; a gateway can stop it happening. The native controls plus OpenTelemetry give you an inventory and an after-the-fact record, which is most of an audit story and none of an enforcement story.

What the standard is doing about it

There is a stable extension for the auth half of this: Enterprise-Managed Authorization (io.modelcontextprotocol/enterprise-managed-authorization). The organisation’s IdP becomes the authoritative decision-maker; the client exchanges an identity assertion for an ID-JAG (Identity Assertion JWT Authorization Grant), then exchanges that for an access token from the server’s authorization server. Onboarding, revocation and policy all move to the IdP. It is the right design.

And now the part that decides whether you can use it

On the official, community-maintained extension support matrix, exactly one client is listed as supporting it: Archestra.AI. Not Claude Desktop, not Claude on the web, not VS Code GitHub Copilot, not Microsoft 365 Copilot, not ChatGPT, not Cursor, not Goose. “Extensions are opt-in and never active by default,” and the docs note that support “typically requires client-level support from the organization’s IT team in addition to the MCP client application.”

So the honest answer to “is the standard solving this?” is: the specification has, and the clients your organisation actually deploys have not. That gap is the entire commercial case for gateways. It is also why the gap will close — which should affect how much you build.

One correction worth making, because the vendor write-ups repeat it: the MCP roadmap (last updated 22 August 2026) does not list gateways or proxies as a priority area. Its five priorities are agentic messaging primitives, HTTP-native transport unification, agent identity and enterprise-ready security, improved primitives, and SDK developer experience. The enterprise work is happening under identity — DPoP, workload identity federation, token exchange — not under a gateway pattern. If someone tells you the protocol is standardising gateways, they have read a blog post rather than the roadmap.

What actually exists

This is the question that separates a decision from a discussion. Figures read from the GitHub API on 14 September 2026.

ImplementationLicenceStarsLast pushWhat it is
IBM ContextForge
IBM/mcp-context-forge
Apache-2.04,47214 Sep 2026The most substantial open implementation. A gateway, registry and proxy in front of MCP, A2A and REST/gRPC, with federation, plugins and OpenTelemetry tracing
Docker MCP Gateway
docker/mcp-gateway
MIT1,56426 Aug 2026A docker mcp CLI plugin. Containerised server execution with a gateway in front — the closest thing to pinning, because the unit you pin is an image
Lasso MCP Gateway
lasso-security/mcp-gateway
MIT38722 Jan 2026Plugin-based orchestration of other MCP servers. Note the date: roughly eight months without a push, in a protocol that revised its specification in that window

There is also a live commercial category — vendors selling managed MCP gateways with SOC 2 positioning — and cloud API-management products adding MCP support. I have deliberately not ranked those, because at this stage the vendor comparisons are mostly written by the vendors.

Where this is still hand-rolled, stated plainly

Three pieces of primary evidence, all :

  • The official MCP Registry is in preview — “Breaking changes or data resets may occur before general availability”
  • It “does not support private servers”, and tells you to host your own private registry if you need them
  • “The official MCP Registry codebase is not designed for self-hosting, and the registry maintainers cannot provide support for this use case” — fork it and you own it

And it does not review code: “The MCP Registry delegates security scanning to” the underlying package registries and downstream aggregators. Its namespace verification proves a server came from the domain it claims. That is authenticity of name, not safety of behaviour.

So: buying a gateway is real and the products are real. An internal catalogue of approved servers, the review process that admits things to it, and anything resembling version pinning are still yours to build, whichever product you buy.

Two things called “gateway” that are not the same thing

Anthropic documents an LLM gateway for Claude Code: you point ANTHROPIC_BASE_URL at it, it holds the provider credential, issues one key per developer, and requests arrive at POST /v1/messages. That buys per-developer attribution, spend control, offboarding and routing. It sits on the inference path and knows nothing about which MCP servers are connected or what tools were called. An MCP gateway sits on the tool-call path and knows nothing about your token spend. Two different chokepoints, one word. Conflating them is the most common way these conversations go wrong, and noticing the difference out loud is a cheap way to be the most precise person in the room.

The specification constrains your gateway

A gateway is an MCP server to the client and an MCP client to the upstream server, which means the authorization security requirements land on it directly. Three normative lines from the 2026-07-28 specification:

  • “If the MCP server makes requests to upstream APIs… The MCP server MUST NOT pass through the token it received from the MCP client.” A gateway that forwards the caller’s token upstream is non-conformant, and it is the obvious first implementation.
  • “MCP servers MUST only accept tokens specifically intended for themselves and MUST reject tokens that do not include them in the audience claim.”
  • “MCP proxy servers using static client IDs MUST obtain user consent for each dynamically registered client before forwarding to third-party authorization servers.” The spec names the proxy case explicitly, because the confused-deputy failure is a proxy failure.

Read together, these say something useful: the protocol anticipated the gateway pattern and constrained it before the products arrived. The gateway is not a hole in the standard. It is a component the standard has security requirements for, which is a much better position to buy from — and a conformance question you can put to a vendor.

The case against, which is not weak

Objection — and it is the real cost of this decision

A gateway reintroduces the centralisation MCP was designed to avoid

MCP’s appeal was decentralised, per-developer extension: an engineer adds a server in a minute and gets leverage without asking anyone. A gateway puts a team, a queue and a change-approval process in front of that minute. It also puts a single process in the path of every tool call in the company. When it is down, nobody’s agent can do anything; when it is slow, every workflow is slow; and when it lags the protocol, everyone is on the old protocol.

What bears on it: This is not speculative, and the evidence is in Anthropic’s own LLM-gateway documentation — the failure modes of a central component on the inference path are documented, and an MCP gateway inherits their shape. Four of them:
  • Buffer instead of streaming and you stall clients: “the whole response arriving at once after a pause means the gateway is buffering”
  • Allowlist headers instead of forwarding them and new capabilities break on client upgrade — the guidance is to “forward anthropic-* headers and request bodies verbatim rather than allowlisting”
  • Rewrite upstream errors and automatic recovery stops working, because the client “matches on error wording”
  • A WAF in front of the gateway “returns 403 on real sessions while short test requests pass” — a central component failing in a way your smoke test cannot see
And the version-lag problem is concrete rather than theoretical: when a new model ships, developers selecting it get a 404 until someone updates the gateway’s routing table. Substitute “new tool” for “new model” and you have the MCP version of the same bottleneck. I think a gateway is the right answer for most organisations above a certain size and that it is not a pure win, and anyone selling it to you as a pure win has not operated one.

So when is it warranted?

Same rule as every other mechanism in this course — the cheapest control that does the job, and the job is defined by blast radius, not by tidiness:

SituationCheapest control that does the job
Read-only servers over public dataNO GATEWAY A denylist and an inventory. You are buying infrastructure to solve a problem you do not have
You need to know what is connectedNO GATEWAY OTEL_LOG_TOOL_DETAILS=1 and aggregate. Inventory is the cheapest audit and it is already there
You need a fixed approved setNO GATEWAY managed-mcp.json, or allowedMcpServers with allowManagedMcpServersOnly: true
Servers hold credentials to systems of recordGATEWAY Credential custody is the job, and it is the one no client-side setting can do
Tool calls write to production, or touch regulated dataGATEWAY You need enforcement in the path, not a record after the fact
Many clients, not just Claude CodeGATEWAY Per-client controls give you N policies and N audit formats. A gateway is the only place one policy can live
Takeaways
  • MCP is the one mechanism where the decider sits outside your organisation — that is what creates the control problem
  • Native controls do allowlisting and provisioning; they do not do credential custody, rate limiting or version pinning
  • Nothing overrides a denylist match; an allowlist without allowManagedMcpServersOnly is user-broadenable; unset ≠ empty
  • The standards answer (Enterprise-Managed Authorization) is stable but has one listed client
  • Real implementations exist — IBM ContextForge, Docker MCP Gateway — but the approved catalogue, the review process and pinning are still yours
  • The spec already constrains gateways: no token passthrough, audience validation, consent on proxy registration
  • A gateway is a bottleneck and a single point of failure. Buy it for credential custody and production blast radius, not for tidiness
7

When the documented cost isn’t the real cost

A case where the architecture is right, the docs are clear, and the measurement disagrees
By the end of this module you will
  • Know what Skills’ progressive disclosure promises and what it costs in practice
  • Be able to compute your own organisation’s exposure rather than trusting a per-unit figure

What the docs promise

Agent Skills use progressive disclosure across three levels, with a published cost per level:

LevelWhen loadedDocumented token cost
1 — MetadataAlways, at startup~100 tokens per Skill (name + description)
2 — InstructionsWhen triggeredUnder 5K tokens (SKILL.md body)
3 — ResourcesAs neededNone until accessed

The promise is explicit: “This lightweight approach means you can install many Skills without context penalty: until a Skill is triggered, only its name and description occupy context.” The architecture is genuinely good, and the filesystem design — scripts run via bash so their code never enters context — is the right idea.

What people actually measure

Issue #14882 — open, labelled bug and has repro

Opened 20 December 2025, still open, 19 comments. Title: “Skills consume full token count at startup instead of progressive disclosure (frontmatter only)”. The reporter shows /context attributing 3–5.5K tokens to each skill from the official plugin-dev plugin.

The thread sharpens as more people measure their own installs:

  • One commenter with ~50 skills reports the full catalogue — every name plus its full description — being re-injected as a system reminder after every tool result and every user turn. That is a per-turn cost, not a startup cost.
  • Another, with 148 SKILL.md files, measures ~41KB of descriptions injecting at session start, about 10K tokens before any user turn, and frames the ticket as a documentation-contract violation rather than a bug in the architecture.
  • A third reports that removing skills does not reduce context: the tokens “move to the System tools row in an exact 1:1 trade.”

Compute your own number — a worked example

The lesson is not “skills are expensive.” It is that a per-unit figure from a doc is not a budget. Here is the measurement, run against a real 52-skill installation:

SKILL.md files installed52
Total name + description characters6,351
→ metadata floor, the progressive-disclosure promise~1,590 tok
Total body characters (all 52)131,975
→ if bodies loaded at startup, the failure mode~33,000 tok
Spread between best and worst case~21×

Note the two things this shows. First, this installation’s descriptions average ~31 tokens, not the ~100 the docs use — the doc figure is a conservative rule of thumb, and your real floor may be lower. Second, the exposure if progressive disclosure does not hold is 21× the floor, and that is the number worth knowing before an org standardises on a large skill library.

The transferable practice

Measure the mechanism in your own installation before you budget for it. Count the metadata, count the bodies, and know the ratio. Then check /context against the floor. If they disagree, you have found your number and the doc has not.

This generalises well beyond skills. Any mechanism with a “loads only when needed” story has a best case and a worst case, and the gap between them is the thing to quantify — because that gap is what you are exposed to when the optimisation does not fire.

Takeaways
  • Documented: ~100 tok/skill at startup, body on trigger, resources on demand
  • Measured by users in an open bug ticket: full descriptions re-injected per turn; ~10K tokens at 148 skills
  • Measured on a real 52-skill install: ~1,590 tok floor vs ~33,000 tok ceiling — 21×
  • Practice: measure the mechanism in your own installation; quantify the best-to-worst gap, because that gap is your exposure
8

Ownership, review and mutual exclusivity

The part generic courses skip, and the part that actually blocks deployments
By the end of this module you will
  • Know which compliance postures are structurally incompatible with which features
  • Have a concrete answer to “who owns the CLAUDE.md” and “how does a skill get approved”

The exclusions that are not negotiable

Anthropic offers two distinct arrangements — zero data retention (ZDR) and HIPAA readiness — and the docs are explicit that they are different things for different purposes: HIPAA readiness “applies a broader set of privacy and security safeguards than ZDR (encryption, access controls, and audit logging) rather than requiring immediate deletion. If your organization handles PHI, HIPAA readiness is the arrangement to use; you do not also need ZDR.”

If your posture is…You cannot have…
Zero data retention EXCLUDED Claude Fable 5.1, Mythos 5.1, Fable 5, Mythos 5 — these “require 30-day data retention and are not available under ZDR unless expressly authorized by Anthropic.” Your ZDR posture excludes the most capable models.
EXCLUDED Agent Skills — “not covered by ZDR arrangements. Skill definitions and execution data are retained according to Anthropic’s standard data retention policy.”
EXCLUDED Claude Managed Agents — stateful; transcripts persist until deleted
EXCLUDED Claude Console and playground, Claude Teams and Enterprise interfaces, Claude for Excel, consumer plans
EXCLUDED CORS — unsupported for ZDR orgs; browser apps must route through a backend proxy. An architecture constraint, not a policy one
HIPAA readiness (BAA) EXCLUDED Claude Code — “Claude Code is not covered under HIPAA readiness.” Flat, no configuration
EXCLUDED Claude Platform on AWS and Microsoft Foundry
EXCLUDED Beta features generally, unless explicitly listed as eligible
EXCLUDED PHI in JSON schemas — structured-output and strict: true schemas are compiled to grammars cached separately and “do not receive the same PHI protections.” Applies to property names, enums, consts and regex patterns
Two things worth carrying into a procurement meeting

1. The API enforces this, so it is a runtime failure, not a policy discussion. A HIPAA-enabled org sending a non-eligible feature gets a 400: “The requested features are not available for HIPAA-regulated organizations without Zero Data Retention: code_execution.” Your compliance posture will show up as broken builds.

2. The sharpest exclusion is the model one. An organisation that adopts ZDR as a blanket default has, without necessarily realising it, excluded itself from Anthropic’s most capable models and from Agent Skills. That is a legitimate tradeoff and it should be made deliberately, by someone who knows they are making it — not inherited from a security questionnaire that had a “zero retention” checkbox.

Who owns what

ArtifactNatural ownerThe failure if unowned
Org-wide context (managed settings, org-tier policy)Platform / securityNobody can enforce anything; every control is advisory
Repo CLAUDE.mdThe repo’s owning team, reviewed like codeGrows additively forever; billed on every turn; no deletion process. This is the default outcome
Shared skillsA named reviewer, plus a distribution channelDuplication, drift, and a skill nobody audited running with your permissions
Retrieval indexData platformResidency and freshness become nobody’s problem
Memory storesWhoever owns the data classificationA persistent store of model-selected content with no retention policy
How a skill should get reviewed before an org depends on it

Anthropic’s own guidance is blunt: use Skills only from trusted sources, because “a malicious Skill can direct Claude to invoke tools or execute code in ways that don’t match the Skill’s stated purpose,” and it says to treat installing one like installing software.

So the review is a software review, not a document review, and it has four questions a reviewer must be able to answer:

  1. What does the description claim, and does the body do only that? The description is what the model matches on — a description broader than the body is a misfire generator, and a body broader than the description is the security case.
  2. What does it execute, and what does it reach? Bundled scripts run with the session’s permissions. External fetches are the highest-risk pattern, because fetched content can carry instructions.
  3. What does it cost? Description length is paid always; body length when triggered. Both are reviewable numbers.
  4. Who is on the hook when it is wrong? A named owner, not a team alias.

Two practical constraints on distribution: custom Skills do not sync across surfaces — claude.ai, the API and Claude Code are three separate uploads — and on claude.ai they are per-user with no org-wide admin management. An org standardising on skills is standardising three times, and cannot centrally manage the claude.ai copy at all. Claude Enterprise organisations can turn on Skill content scanning for claude.ai and Cowork uploads, but that scanning does not cover the Skills API or the Console.

Takeaways
  • ZDR excludes Fable 5.1 / Mythos 5.1 / Fable 5 / Mythos 5, Agent Skills, Managed Agents, Console, and CORS
  • HIPAA readiness excludes Claude Code entirely, Claude Platform on AWS, Foundry, and beta features — and PHI must stay out of JSON schemas
  • The API returns a 400 naming the ineligible feature — compliance surfaces as broken builds
  • Skill review is a software review: description-vs-body, execution and reach, cost, named owner
  • Skills do not sync across surfaces, and claude.ai skills have no org-wide admin management
9

Practices that earn their place

Organised by the axis each one protects, because a practice with no failure mode is decoration
By the end of this module you will
  • Have a short set of practices you can defend one at a time
  • Be able to say, for each, what breaks without it
  • Know which commonly-recommended practices this course declines to recommend, and why

The test for whether a practice belongs in your rollout is not whether it sounds responsible. It is whether you can finish this sentence: “if we skip this, here is the specific thing that goes wrong, and here is how we would find out.” A practice that cannot finish that sentence is asking engineers to pay a tax so that a document can exist.

So these are grouped by the axis each one protects, rather than listed. If you only take three, take the first one from each group.

Protecting the decision-time axis

PracticeWhy, and what breaks without it
Write down the rate of change before choosing the mechanism One line in the PR: this content changes <never / quarterly / daily / per-request>. It converts an argument about tools into a question of fact.
Without it: the quarterly pricing table lands in a CLAUDE.md, because “quarterly” sounds stable. It is then stale for most of each quarter and billed on every turn of every session that never needed it. This is the single most common mechanism error and it is invisible — nothing fails, it is just wrong and expensive.
Put a byte-stability rule on everything above the cache breakpoint No timestamps, no request IDs, no “Hello, {name}” in the system prompt. Caching is a prefix match; any byte change invalidates everything after it.
Without it: a one-line, well-intentioned edit takes your organisation’s cache hit rate to zero with no error and no alert, and the bill rises roughly tenfold on the input side. See Module 2.
Re-decide on a cadence, not on complaint Set a date to revisit each artifact’s placement. Rates of change themselves change.
Without it: placement decisions are only ever revisited when someone is annoyed, which selects for the loud problems and never touches the expensive quiet ones.

Protecting the decider axis

PracticeWhy, and what breaks without it
Name an owner and a deletion cadence for every context artifact Ownership is the easy half and most orgs do it. The deletion cadence is the half that is always missing.
Without it: a CLAUDE.md grows additively forever, because every addition has an advocate and no removal has one. It is the only mechanism billed unconditionally on every turn, so an unowned one is a permanent, silent, compounding line item.
Gate on distribution, not on authoring Anyone may write a skill or a CLAUDE.md for their own repo with no process at all. Review attaches at the moment other people start depending on it.
Without it: you either gate everything — and turn a two-hour improvement into a two-week one, which is the objection in Module 12 and it is correct — or you gate nothing, and an unreviewed artifact ends up running with the whole organisation’s permissions.
Prefer the denylist for anything that must never load “Nothing overrides a denylist match,” while allowlists merge across settings scopes unless allowManagedMcpServersOnly is set.
Without it: you have an allowlist in configuration management that a developer can widen in their own settings file, and an enforcement posture that exists only in the document describing it.

Protecting the per-turn-cost axis

PracticeWhy, and what breaks without it
Instrument cache_read_input_tokens before you optimise anything else It is one field in usage, and it is the difference between paying full input rate and 10% of it on every repeated turn.
Without it: you cannot tell a working cache from a broken one, because a broken cache raises no error. Teams have run for months believing caching was on. This is the cheapest instrumentation in the whole stack and the highest-leverage.
Budget skill descriptions, not just skill bodies Descriptions are the always-paid floor; bodies are conditional. On a 52-skill installation the spread between the two was 21× — which means the review question “is this skill worth it?” has two different answers depending on which number you meant.
Without it: the unconditional floor grows with every skill anyone adds, nobody attributes the cost to the skills, and the measured-versus-documented dispute in Module 7 decides your bill instead of you.
Route latency-tolerant work to Batch, and decide model tier only after caching Batch is 50% off and the docs state the multipliers stack with other modifiers. Caching and batch are free levers; model choice is a quality tradeoff.
Without it: the first cost conversation becomes “should we downgrade the model,” which spends quality to buy savings you could have had for nothing.
One practice that spans all three axes, and belongs to you specifically

Make your enforcement posture version-aware. Claude Code’s v2.1.259 changed how allowedMcpServers applies to managed servers, and servers an admin had suppressed that way “start loading on each user’s first launch of v2.1.259 or later, with no prompt or notice.” On a client that releases near-daily, a security control you validated once is a security control you validated against one version. The practice is to test new client releases against your managed configuration before they reach the fleet — which is the same advice Anthropic gives for LLM gateways, for the same reason.

Four practices I am not recommending, and why

Naming these matters more than adding a tenth practice, because each one will be proposed to you and each one sounds obviously correct.

Commonly proposedWhy it doesn’t earn its place
A central prompt libraryPrompts are the least stable artifact in the system and the least reusable across codebases. A library of them decays faster than anyone maintains it, and the thing that actually transfers — a working technique demonstrated on your own repository — transfers better as a message in a channel than as a catalogue entry.
A per-team token budgetIt looks like cost control and behaves like a disincentive to use the tool, which is the opposite of what you were hired to do. Budget the unconditional costs — the prefix, the skill descriptions, the always-loaded context — because those are the ones nobody is choosing. Metered usage that tracks work being done is not the problem.
An AI centre of excellenceIt centralises the expertise at exactly the moment you need it distributed, and it creates a queue. Anthropic’s own adoption guidance points the other way: the signal that adoption is working is “questions in the channel are being answered by people other than you.” A team whose job is to answer the questions guarantees that never happens.
Mandatory training before accessIt front-loads the cost and delays the only thing that reliably converts a sceptic, which is one successful result on their own code. Keep the material, drop the gate.
Takeaways
  • A practice earns its place only if you can name what breaks without it
  • Decision time: write down the rate of change; byte-stability above the breakpoint; re-decide on a cadence
  • Decider: owner plus deletion cadence; gate on distribution not authoring; prefer the denylist
  • Per-turn cost: instrument cache reads first; budget skill descriptions; batch before downgrading the model
  • Enforcement has a version number on a near-daily-release client
  • Declined: prompt library, per-team token budget, centre of excellence, training-before-access
10

Driving adoption

Sequencing, the stall point, and how to tell real adoption from the appearance of it
By the end of this module you will
  • Have a defensible order of operations, and know why each step comes where it does
  • Know the one place adoption efforts predictably stall, from the vendor’s own guidance
  • Know what the available metrics measure and, more usefully, what they cannot
  • Know which arguments you will meet and which of them are correct

The finding that should set your expectations

DORA’s position on AI-assisted development is that “AI’s primary role is as an amplifier, magnifying an organization’s existing strengths and weaknesses,” and that “the greatest returns on AI investment come not from the tools themselves, but from a strategic focus on the underlying organizational system.”

DORA’s follow-up ROI report frames returns as following a J-curve — an initial dip before value — with the dip attributed to the extra effort of verifying AI-generated work. I have read the framing on DORA’s own pages but not the underlying figures, so treat the shape as reported and the magnitude as unquoted.

Two consequences for someone arriving to drive adoption. First, if the surrounding system is weak — slow review, unclear ownership, a flaky test suite — the tool will make that worse before it makes anything better, and you should say so in advance rather than be surprised by it in month two. Second, a leader who expects a straight line up will interpret the normal early dip as failure. Predicting the dip out loud, before it happens, is the cheapest credibility you will ever buy.

Sequencing: what to do first, and why that order

The ordering principle is simple: do the steps whose outcome would invalidate later work first. Most rollouts run these in exactly the wrong order, because the visible steps are the fun ones.

#StepWhy here, and what it would have cost you later
0Establish the compliance posture before the pilot, not after it Module 8’s exclusions are structural, not settings. A ZDR default forecloses the frontier models and Agent Skills; HIPAA readiness excludes Claude Code entirely. Discovering this after a successful pilot is the single most expensive sequencing error available, because you lose the pilot and the credibility that came with it.
1One team, real work, a bounded window Not a bake-off, not a survey. A team whose work you can actually see, doing work that would have happened anyway. The output you need is a decision, and a pilot that cannot produce one is a demonstration.
2Instrument the cheap things immediately cache_read_input_tokens and, if you are on Claude Code, OpenTelemetry with OTEL_LOG_TOOL_DETAILS=1. Both are near-free and both answer questions you will be asked in month three. Retrofitting instrumentation means your first three months have no baseline.
3Champions, not a programme See below. This is where adoption is won, and it is the step most likely to be replaced by a training plan.
4Name owners for context artifacts while there are still few of them Retrofitting ownership onto forty CLAUDE.md files is a different and much worse job than establishing it over four.
5Policy and gateways last, sized to actual blast radius You cannot size a control before you know what people actually connect to — which is what step 2 tells you. Building the gateway first means building it for the usage you imagined.

The stall point, named

Where adoption efforts predictably die

Anthropic’s own adoption guidance puts it in one sentence: a concrete, runnable example “removes the gap between curiosity and a first successful use, which is where most adoption efforts stall.” And on the mechanism: “Adoption of a developer tool rarely happens because of a rollout announcement. It happens because someone on the team begins using the tool well, talks about it openly, and makes it easy for others to follow.”

The guidance is specific about what that costs: roughly 15 minutes a week posting wins and prompts, 20 minutes answering questions in a shared channel, 5 minutes running a weekly show-and-tell thread, and 0–30 minutes of optional pairing. That budget is the argument. You are not asking for headcount; you are asking for forty minutes a week from people who are already enthusiastic. A proposal priced that way is very hard to refuse, and a proposal that asks for a programme is very easy to defer.

One organisational point the same guidance makes, which is worth defending: on “what about security and data handling?” it says to refer the question to the administrator, because “champions should not improvise this answer.” That is a correct separation of duties, and it also protects your champions — the fastest way to lose one is to let them answer a compliance question wrongly in public.

Real, or theatre?

Three tests. The first is borrowed and is the best of them.

TestRealTheatre
Who answers the questions? — this is the vendor’s own week-four success signal: “questions in the channel are being answered by people other than you”Several people answer; threads resolve without youEvery thread waits for one person, or for a vendor
Return versus trialThe same names appear week after weekA large, flat trial count with no repeat cohort — which a “% of engineers who have used it” metric reports as success
Does it appear in work that ships?Usage shows up in merged pull requests and in incident timelinesUsage shows up in a dashboard and a quarterly slide

What to measure, and what each number cannot tell you

The Claude Code analytics dashboard exposes a specific set, and the limits are documented as precisely as the metrics.

MetricWhat it actually saysWhat it cannot tell you
Lines of code acceptedLines the user accepted in-sessionWhether any of it survived. The docs state it “excludes rejected suggestions and does not track subsequent deletions”
Suggestion accept rateHow often edit-tool suggestions are acceptedWhether accepting was correct. A high rate is equally consistent with good output and with insufficient review
PRs with CC / Lines with CCMerged PRs containing attributed linesCausation. Attribution matches sessions to PR diffs in a window of 21 days before to 2 days after merge, and code “substantially rewritten by developers, with more than 20% difference, is not attributed”
Daily active users / sessionsEngagementValue. This is the number most likely to be presented as a result
LeaderboardTop contributors by volumeAlmost nothing you want. Its documented use is finding people who can help others; the moment it is read as a performance ranking it manufactures exactly the usage it claims to measure

The docs say the contribution metrics are “deliberately conservative and represent an underestimate… Only lines and PRs where there is high confidence in Claude Code’s involvement are counted.” Say that out loud when you present the numbers. A leader who learns later that a metric was conservative discounts everything you showed them; a leader who is told in advance treats it as a floor.

The exclusion that will catch you, and it is new to this course

“Contribution metrics are not available for organizations with Zero Data Retention enabled. The analytics dashboard will show usage metrics only.”

Read that against Module 8. A ZDR posture already costs you the frontier models and Agent Skills. It also costs you the PR-level attribution — which is to say, the compliance posture adopted to get the deployment approved removes the evidence you need to prove the deployment worked. You keep daily active users and lines accepted; you lose the link to shipped work.

This is the same structure as every other exclusion in this course and it is the one most likely to be discovered at the worst moment, because nobody puts “how will we measure this” and “what is our retention posture” in the same meeting. If your organisation is heading for ZDR, agree the alternative evidence — cycle time, review latency, a before-and-after on a named workflow — before it is switched on, while you can still capture a baseline.

The arguments you will meet

Which of these is correct matters more than having a response to each. Three of the six are right, and conceding them quickly is what makes the other three land.

ArgumentIs it right?The honest reply
“I’m faster without it.”OFTEN RIGHTThe vendor’s own guidance concedes this: it is “likely true for code the person writes routinely.” The reply is a redirect, not a rebuttal — suggest the work they avoid: legacy files, unfamiliar services, test scaffolding.
“Review will become the bottleneck.”RIGHTGeneration got cheap and review did not. This is the amplifier finding applied to one queue. If your review capacity is already the constraint, adoption makes it worse first — plan for review, not just for generation.
“Governance will slow us down.”PARTLY RIGHTRight for personal and team-scoped artifacts, which should have no gate at all. Wrong where blast radius is organisational. Gate on distribution, not authoring (Module 9).
“We tried it, it hallucinated.”USUALLY NOTUsually a context problem rather than a model problem — the whole of this course is about what was in the window. Re-run the same prompt with the relevant files referenced and the actual error output pasted in.
“The window is 1M now, so context management is obsolete.”NOCapacity was one of five constraints and the only one that got solved. Filling 1M tokens on Opus 5 is about $5 a turn uncached (Module 3).
“Let’s wait for the standards to settle.”NO, BUT…It is a reasonable instinct pointed at the wrong layer. Waiting on mechanism choices is sensible — those are vendor-shaped and will be renamed. Waiting on posture, instrumentation and ownership is not, because those are the things that take months and gate everything else.
Takeaways
  • AI is an amplifier of the existing system; predict the early dip out loud before it happens
  • Sequence by what would invalidate later work: posture → one real team → instrument → champions → owners → policy
  • The documented stall point is the gap between curiosity and a first successful use, and it costs ~40 min/week to close
  • Week-four success signal: questions answered by people other than you
  • Acceptance metrics do not track deletions; attribution drops code rewritten by >20%
  • ZDR removes contribution metrics — agree alternative evidence and take a baseline first
  • Of the six arguments you will meet, two are right and one is partly right. Concede those fast
11

Three decisions, scored

Mechanism selection, the economics, and the room
By the end of this module you will
  • Have applied the rate-of-change rule to ten real situations
  • Have tested your intuitions about where an agentic bill concentrates
  • Have practised the harder skill: judging which objections are correct before answering them

Ten situations. For each, pick the cheapest decision moment that still catches the rate of change — not the most powerful mechanism available.

Tool 1

Cheapest mechanism that catches the change

0 / 10
Result

Tool 2 — where does the money go?

Eight claims about token economics. True or false, on the figures in this course.

Tool 2

Economics: true or false

0 / 8
Result

Tool 3 — the room

Ten things that will be said to you, by people who are not being difficult. For each one, decide whether they are right, partly right, or wrong — before you decide what to say. This is the tool that matters most. Conceding the correct objections quickly is what buys you the standing to push back on the others, and an adoption lead who argues all ten loses the room in the first meeting.

Tool 3

Right, partly right, or wrong?

0 / 10
Result

Takeaways
  • The rule is rate of change → decision moment, not capability → mechanism
  • The mechanisms people over-reach for are retrieval and MCP; the one they under-price is CLAUDE.md
  • After caching, the bill is dominated by output and turn count
  • In the room, three of the ten objections are right and three are partly right — concede those first
12

Counterarguments

Marked as objections. Four of the seven are aimed at this course.
By the end of this module you will
  • Know the strongest cases against organising your thinking this way
  • Know which of them I think is right
Objection 1 — and I think this one lands

Most enterprise agent deployments fail on process, not on any of this

The failures that actually kill deployments are: no owner, no eval set, no definition of done, no rollback, unclear escalation, and a pilot that was never scoped to produce a decision. None of those are context-window problems. An organisation can get every mechanism choice in this course right and still fail, and an organisation can get all of them wrong and succeed because a competent team iterated.

Worse, this material is attractive to the wrong audience: it is tractable, quantitative and feels like progress, which makes it an excellent way to avoid the harder organisational questions. A token-economics review is a more comfortable meeting than “who is accountable when this is wrong.”

What bears on it: I think this is correct and it is the most important objection in the course. My defence is narrow: this material is necessary and nowhere near sufficient. It matters at exactly two moments — when you are choosing a mechanism (Module 5) and when procurement hits a compliance wall (Module 8) — and it is close to irrelevant the rest of the time. If you are choosing between reading this and writing an eval set, write the eval set.
Objection 2 — also aimed at this course

The mechanism taxonomy is vendor-shaped, not problem-shaped

Skills, CLAUDE.md, MCP and subagents are not natural categories. They are one vendor’s product decisions, several of which did not exist eighteen months ago and several of which will be renamed or absorbed. Building an organisation’s mental model around them is building on a roadmap. The problem-shaped version is far shorter: you have a budget of tokens per turn; decide what goes in it and who decides. Everything else is this quarter’s packaging.

What bears on it: Largely right, and it is why the course is organised around three axes rather than six product names — decision time, decider and per-turn cost are properties of any system that assembles a context window, including a competitor’s and including one you build. The axes should survive the renames. But the objection still bites where I named specific products, and a reader should treat Module 5’s rows as examples of the axes rather than as the taxonomy.
Objection 3 — aimed at Module 6

A gateway is an enterprise reflex, and it will be obsolete before it pays for itself

The entire commercial case for MCP gateways rests on a gap: the standards-track answer exists and the clients don’t implement it. But that gap is the most predictable thing in this course. There is a stable Enterprise-Managed Authorization extension, a Working Group on agent identity, and every major client vendor has an enterprise sales motion that rewards implementing it. An organisation that spends two quarters building an MCP gateway is building a bridge over a river that is being drained, and will then own it — with its on-call rota, its version lag and its outage — for years after the thing it existed to work around has shipped natively.

What bears on it: Half right, and the half that is right should change what you build rather than whether you build. Credential custody is not the part that gets solved by client support for EMA: an IdP deciding who may reach a server is a different problem from where the upstream system’s credential lives and who can read it. That half is durable. Allowlisting and central auth are the parts most likely to arrive natively — Claude Code already does the allowlisting half — so building those yourself is the reflex the objection correctly identifies. The practical form: buy or run something small for custody and enforcement on the paths with real blast radius, and let the catalogue and the auth flow arrive.
Objection 4 — aimed at Module 10

Every metric in the measurement section is an activity metric, so measuring adoption at all is theatre

Look at what is actually available: daily active users, sessions, lines accepted, acceptance rate, attributed lines in merged PRs. Not one of those measures whether the software got better, shipped sooner, or broke less. Even the best of them — PR attribution — is a measure of how much model-produced text survived into a diff, which is closer to a volume metric than a value metric, and it quietly rewards the behaviour it should be suspicious of. Presenting these to a leadership team as evidence of value is the theatre the module claims to be detecting.

What bears on it: This one is uncomfortable and largely correct, and I have not solved it. The defence is only that activity metrics are useful for the narrow thing they measure — whether people came back — and actively misleading for everything else, which is why the module pairs each metric with what it cannot say and puts the three qualitative tests first. The honest position is that the value question is answered by the same instruments you already trust for engineering outcomes — cycle time, review latency, change failure rate, on a named workflow with a before and after — and that the AI dashboard is an adoption instrument, not an outcome instrument. If someone tells you those two are the same thing, they are selling something, and it may be me.
Objection 5

Optimising token spend is optimising the wrong denominator

Model costs have fallen persistently and engineer time has not. A context optimisation that saves $12,500 a year and costs two engineer-weeks a year to maintain is a loss. Worse, per-turn cost is the metric that is easy to measure, so it attracts attention out of proportion to its size — the right denominator is cost per completed task, which is often dominated by retries, wrong turns and human review.

What bears on it: Right, and the course’s own arithmetic supports it: the recommended ordering puts the two free levers (caching, batch) first precisely because they cost no engineer time. Everything past those is a tradeoff. The one place I would defend spend-watching is that agentic loops fail toward cost — a broken loop burns budget silently, so the number is a fault detector even when it is not a lever.
Objection 6

The compliance exclusions are a snapshot, not a structure

Module 8 reads as though ZDR-excludes-frontier-models is a law. It is a current product decision that could change with one contract negotiation — the docs even say “unless expressly authorized by Anthropic” and tell you to check your contract terms. An enterprise that hard-codes today’s matrix into its architecture will be wrong in a year.

What bears on it: Fair, and the “checked 14 September 2026” stamp is there for this reason. The durable part is not the matrix but the shape: statefulness and zero retention are in structural tension, and that tension will outlive any specific exclusion list. The right posture is to know which of your requirements are structural and which are this quarter’s policy, and to re-check the latter at renewal.
Objection 7

Centralising context governance will just slow everything down

Review boards for CLAUDE.md files and approval gates for skills are how organisations turn a two-hour improvement into a two-week one. The teams getting value from these tools are the ones iterating fastest, and every control proposed in Module 8 is friction applied to exactly that.

What bears on it: The strongest form of this is right for personal and team-scoped artifacts, and I would not gate those at all. It is wrong for the two cases where the blast radius is organisational: a skill distributed to everyone, and a context artifact that reaches production output. The distinction is not centralisation versus speed, it is scope — gate on distribution, not on authoring. Anyone can write a skill; the gate is the moment it becomes something other people depend on.

Three that do not survive the material

The objectionWhy it fails
“Context management is obsolete now that windows are 1M.” Capacity was one of five constraints and the only one that got solved. Cost, latency, attention and governance all got harder, and cost now scales with adoption. Module 3.
“Just cache everything and the cost problem goes away.” Caching is a prefix match that does nothing for output, and output was 38% of the modelled bill after caching. It is the first lever, not the only one. Module 2.
“Put it in the CLAUDE.md” as a default answer. It is the only mechanism billed on every turn of every session unconditionally, and the one with no natural deletion process. It is the right answer for stable, universal, short content and the wrong one for everything else. Module 5.
Where this leaves a decision-maker

Four things are worth doing on Monday, in order, and none of them is a context-window decision:

  1. Turn on prompt caching and verify the hit rate. Free, large, and the verification step is the one people skip.
  2. Find out which compliance posture you have actually committed to, and what it excludes. Most organisations discover this at the 400, not at the contract.
  3. Name an owner for each context artifact, with a deletion cadence. The absence of a deletion process is the failure mode, not the absence of a review process.
  4. Measure your own mechanism costs rather than trusting per-unit figures — and then go and write the eval set, because objection 1 is right.
Takeaways
  • Objection 1 is right: process failures dominate, and this material is necessary but far from sufficient
  • Objection 2 is largely right: the axes are durable, the product names are not
  • Objection 3 is half right: build for credential custody, which is durable; let the catalogue and the auth flow arrive
  • Objection 4 is largely right: the dashboard is an adoption instrument, not an outcome instrument — and I have not solved it
  • The durable compliance insight is statefulness vs zero retention, not the current exclusion list
  • Governance answer: gate on distribution, not authoring
  • Monday list: verify caching → find your real compliance posture → name owners with a deletion cadence → measure, then write the eval set