Skills, CLAUDE.md, memory, retrieval, MCP and subagents are not six features. They are six competing answers to one question — what should the model see right now — and they differ on when the decision gets made, who makes it, and what it costs every turn. Written for the person who has to walk into a room of engineers and leaders and be credible about all of it, then go and actually get it adopted.
Primary sources: models overview · prompt caching · Agent Skills · API and data retention · issue #14882 · managed MCP · Enterprise-Managed Authorization · MCP authorization security · Claude Code analytics · champion kit
This is for someone arriving to help an engineering organisation adopt AI — the person who will be asked, in the same week, what this costs, who owns it, whether it is compliant, and why the pilot has gone quiet. Two things follow. The mechanics are not optional: the token arithmetic and the mechanism taxonomy in Modules 1–8 are what separate someone credible in a room of engineers from someone hand-waving, and hand-waving is detected in about ninety seconds. And the mechanics are not the job: Modules 9–11 are about sequencing, what to do first, how to tell real adoption from the appearance of it, and which of the arguments you will meet are actually correct.
You should finish able to lead a conversation about this rather than survive one — which means knowing which objections to concede in the first thirty seconds so that the rest of what you say lands. It does not explain what a token is beyond what the arithmetic needs.
Prices are as of 14 September 2026. They move, and the arithmetic here is built so you can redo it with whatever the numbers are when you read this.
Short on time? Module 2 for where the money actually goes, Module 8 for the compliance exclusions that are mutually exclusive, and Module 10 for sequencing and what the adoption metrics cannot tell you. Those three change decisions. Module 11 is where you find out whether any of it stuck.
A language model is a function from a context window to a completion. It has no state between calls. Everything an enterprise buys, builds or governs around it — skills, CLAUDE.md files, memory products, retrieval pipelines, MCP servers, subagents, fine-tunes — exists to answer one question:
What should be in the context window at this moment, and what did it cost to put it there?
Treating these as six features to evaluate separately is the single most expensive framing error in enterprise AI adoption, because it produces six procurement conversations, six owners and six budgets for what is one resource-allocation problem.
Most bad mechanism choices are decision-time mismatches: the mechanism’s decision moment does not match how often the answer actually changes.
The rule that falls out: match the decision moment to the rate of change, and match the decider to whoever is accountable when it is wrong. Module 5 applies it mechanism by mechanism; Module 11 makes you do it.
| Model | Input /MTok | Output /MTok | Cache read | Context | Max output |
|---|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $50 | $0.25 (2.5%) | 1M | 128K |
| Claude Opus 5 | $5 | $25 | $0.50 (10%) | 1M | 128K |
| Claude Sonnet 5 | $2 | $10 | $0.20 (10%) | 1M | 128K |
| Claude Haiku 4.5 | $1 | $5 | $0.10 (10%) | 200K | 64K |
From the models overview and pricing pages. Three multipliers stack on top: cache writes cost 1.25× base input at the default 5-minute TTL and 2× at 1-hour TTL; cache reads cost 10% of base input (2.5% on Fable 5.1 and Mythos 5.1); and the Batch API is 50% off. US-only inference is 1.1×. Output is priced at 5× input across the lineup — hold on to that ratio, it drives everything below.
Take a realistic agentic task on Opus 5: a coding agent doing 20 turns, with a 30K-token stable prefix (system prompt, CLAUDE.md, tool definitions) and roughly 2K tokens of new content and 1K tokens of output per turn. — the prices are real, the shape is illustrative.
Now switch on prompt caching for the stable prefix.
Caching cut the bill by 79%. And look at what is left: output is now 38% of the total, up from 8%. It did not get more expensive — everything around it got cheaper.
This is the finding budget models miss. Teams optimise input because input is what they can see growing, and input is the part that caching already fixes for free. Once caching is on, the remaining lever is output volume — how verbose the model is, how many turns the loop takes, how much it re-reasons. Those are controlled by effort settings, prompt design and loop architecture, not by trimming context.
Every turn resends the whole history. A 20-turn loop with a 30K prefix pays for that prefix 20 times. This looks alarming and is almost entirely solved by prompt caching, which is a configuration change, not an architecture change.
The trap: caching is a prefix match — any byte change anywhere in the prefix
invalidates everything after it. A timestamp in the system prompt, a re-ordered JSON key or a varying
tool list silently drops your hit rate to zero, with no error. Check usage.cache_read_input_tokens; if it is zero across repeated requests, something is
invalidating.
Priced at 5× input, never cacheable, and in a loop the output of turn n becomes the input of every turn after it — so a verbose turn is billed once at output rates and then repeatedly at input rates.
That compounding is the real cost of verbosity, and it is why effort settings and loop design matter more to an agentic bill than context trimming. A model that solves the task in 8 turns instead of 20 saves more than any context optimisation available to you.
The practical ordering for an enterprise: turn on caching (free, large), then batch anything latency-tolerant (50%, free), then tune effort per route (measurable), and only then argue about context contents. Most organisations do these in exactly the reverse order because context contents is the part that feels like engineering.
Fable 5.1, Opus 5 and Sonnet 5 all carry 1M-token context windows; Haiku 4.5 carries 200K. The docs put 1M at roughly 555,000 words on the current tokenizer. That is a large fraction of an internal knowledge base.
| Constraint | Does a bigger window relax it? |
|---|---|
| Capacity — will it fit? | YES Genuinely solved for most enterprise documents. This is the constraint people mean, and it is the one that is gone. |
| Cost — what does it cost per turn? | NO Strictly worse. Filling 1M tokens on Opus 5 costs $5 per turn uncached, or $0.50 per turn on cache reads. A 20-turn session on a full window is $10–$100 depending on caching. Capacity became affordable to use long before it became affordable to fill. |
| Attention — will the model actually use it? | NO Relevance density still matters; a fact competing with 900K tokens of other material is not in the same position as one competing with 9K. Long-context benchmarks have improved markedly, but "it fits" and "it is attended to" remain different claims. |
| Latency — how long will it take? | NO More input is more time to first token. For an interactive tool this is a product constraint, not an accounting one. |
| Governance — should this content be there? | NO A bigger window makes it easier to include material nobody approved. Capacity removes the accidental forcing function that used to make someone choose. |
The window got bigger; the budget did not. Context stopped being scarce in the capacity sense and remains scarce in the economic sense — and the economic constraint is the one that scales with your usage rather than sitting fixed. An organisation that treats 1M tokens as “we no longer need to choose” has converted a design constraint into a line item that grows with adoption.
| Context | Retrieval | Memory | |
|---|---|---|---|
| What it is | The tokens in this one request | A lookup that selects tokens to put in this request | State that persists across requests and is written back |
| Lifetime | One call | The index lives; each result is per-call | Indefinite, until deleted |
| Who decides | Whoever assembled the request | Usually the model, via its query | Usually the model, via what it chooses to write |
| Cost shape | Per turn, every turn | Per hit, plus index upkeep | Storage, plus per-turn cost of whatever is loaded back in |
| Compliance surface | Transient — governed by retention terms | Your index; your data residency | A new persistent store of model-selected content |
“Memory” is the word doing the most work in enterprise AI sales, and it is routinely used for all three. Ask one question to disambiguate: “After this conversation ends, what is written down, where does it live, and who chose what went in it?”
Only the third creates a genuinely new compliance object. That is why it matters that the three are priced and procured as if interchangeable.
The retention docs make the practical version of this concrete: Claude Managed Agents — a stateful, memory-shaped product — is called out as not ZDR-covered precisely because “Claude Managed Agents is a stateful resource; session transcripts persist until you delete them.” Statefulness and zero retention are the same tradeoff wearing two names.
| Mechanism | Decides at | Decider | Per-turn cost | Over-engineering when… |
|---|---|---|---|---|
| CLAUDE.md | Authoring time | Whoever owns the repo | Every turn, always. Sits in the cached prefix, so ~10% of input rate after the first write — but never zero, and never conditional | The content is conditional, volatile, or long. Everything in it is a tax on every session including the ones it is irrelevant to |
| Skill | Invocation time | The model, matching your description | Description always; body only when triggered | The procedure is one paragraph, or it always applies (then it is a CLAUDE.md), or it never applies (then delete it) |
| Retrieval / RAG | Turn time | The model, via its query | Per hit, plus index maintenance you own forever | The corpus is small enough to sit in the prefix, or static enough to be a skill. A RAG pipeline for 40 pages of policy is a maintenance obligation bought to solve a formatting problem |
| MCP server | Turn time | The model, choosing a tool | Tool definitions every turn; results per call | The data is static. MCP earns its keep on live external state. Tool definitions are permanent prefix cost — a server exposing 40 tools taxes every turn |
| Subagent | Call time | The model, deciding to delegate | Its own full context, paid once and discarded | The sub-task is short. You are paying a fresh context assembly to save a context you were not close to exhausting |
| Fine-tuning | Training time | Whoever ran the job | Zero per turn — the only mechanism with none | Anything that changes. Zero per-turn cost is bought with maximum change latency, and most enterprise knowledge changes faster than a retrain cycle |
Read the “decides at” and “per-turn cost” columns together and they are the same column. The earlier a mechanism decides, the cheaper it is per turn and the staler it is allowed to get. Fine-tuning decides earliest and costs nothing per turn; retrieval decides latest and costs on every hit. There is no mechanism that is both current and free, and there never will be, because currency is purchased with per-turn work.
So the selection rule is not “which is best” but: how often does this content change, and what is the cheapest decision moment that still catches the change?
Because it is a text file in a repo, a CLAUDE.md feels like documentation, and organisations grow them the way they grow READMEs — additively, with no deletion process. But it is loaded into every session, so every line is billed on every turn of every session by every engineer, forever, whether or not it is relevant.
At ~10% of input rate on cache reads, a 5,000-token CLAUDE.md costs about $0.0025 per turn on Opus 5. Across 200 engineers doing 100 turns a day that is about $50 a day, $12,500 a year, to carry one file. That is not a crisis. It is, though, a recurring cost with no owner and no review cadence, which is exactly the profile of costs that grow unexamined.
Go back to the three axes. MCP is the turn-time mechanism whose decider is the server — the tool list, the tool descriptions and the returned content are all authored by someone outside your organisation and injected into your context at request time. That combination is unique among the six. A CLAUDE.md is yours. A skill is yours. A retrieval index is yours. An MCP server is a third party with a live seat inside every turn.
Anthropic states the position plainly: it “reviews connectors against its listing criteria before adding them to the Anthropic Directory, but doesn’t security-audit or manage any MCP server.” The default is also explicit: “By default, anyone running Claude Code can connect any MCP server they choose.”
| Job | What goes wrong without it |
|---|---|
| Credential custody | Every developer holds a long-lived token for every upstream system on their laptop. Offboarding means finding them. The managed-MCP docs warn directly: “Any user on the machine can read this file, so don’t store API keys or other credentials in env blocks.” |
| A single audit point | You can reconstruct what a model said but not what it did to your systems, and each client keeps its own partial record in its own format. |
| Rate limiting | An agentic loop is a tool-call amplifier. One bad prompt becomes ten thousand calls against a system whose owner never agreed to that load. |
| Allowlisting which servers exist | The set of third parties with a seat in your engineers’ context windows is whatever each engineer chose, and nobody has the list. |
| Version pinning | A remote server changes a tool description overnight and every agent in the company behaves differently the next morning, with no diff, no release note and no rollback. |
The other four have recognisable enterprise analogues — they are secrets management, logging, quotas and software allowlisting, and your organisation already has opinions about all four. Version pinning is different, because the artifact that changes is prose the model reads. A tool description is a prompt. Editing it changes model behaviour across every workflow that touches it, and it passes through none of the review that a code change would. You have a dependency you cannot pin, cannot diff and cannot roll back.
Claude Code ships real administrative control over which MCP servers load. It is worth knowing precisely, because it is more than most people assume and less than a gateway.
| Control | What it does |
|---|---|
managed-mcp.json | Deploys a fixed set. Exclusive: in cloud sessions on a host where the file is deployed, Claude Code “starts with the managed servers only and skips the claude.ai connectors and other servers the cloud host delivers” |
managedMcpServers | Provides org servers alongside what users add. Takes precedence over a same-named local, project, user or plugin server. Requires v2.1.259+ |
allowedMcpServers | Allowlist by serverUrl (* wildcards), serverCommand, or serverName (exact match only, no wildcards) |
deniedMcpServers | Denylist, same entry shapes. “Nothing overrides a denylist match” |
allowManagedMcpServersOnly | Makes the allowlist authoritative. Without it, allowlists from every settings scope merge — including the user’s own ~/.claude/settings.json |
OTEL_LOG_TOOL_DETAILS=1 | Records MCP server and tool names in OpenTelemetry tool events, so you can aggregate which servers users actually connect to |
1. A soft allowlist is user-broadenable. Set allowedMcpServers without allowManagedMcpServersOnly: true and every
settings scope merges, so a developer can widen your allowlist in their own settings file. You will
still see an allowlist in your configuration management and believe it is enforcing something.
2. Unset and empty are opposite. An unset
allowedMcpServers means all servers allowed. Set to [] it means
no servers allowed. The distance between “we haven’t configured that” and
“we configured that to empty” is the whole policy.
3. A client upgrade silently changed enforcement. Before v2.1.259, every managed server had to pass the allowlist. After it, allowedMcpServers
no longer applies to managed-mcp.json servers except where ${VAR} expansion is
used. The docs say so explicitly: servers you had suppressed that way “start loading on each
user’s first launch of v2.1.259 or later, with no prompt or notice.”
Whatever else you take from this module, take this: on a
fast-moving client, your enforcement posture has a version number.
Where it stops. These controls decide which servers load. They are client-side and per-client. They do not hold credentials on your behalf, do not rate limit, do not sit in the path of a tool call to block it, and — the significant gap — there is no version pinning: a server is identified by name, URL or command, never by a version or a digest. And “Claude Code doesn’t have a built-in MCP server registry that users can browse and install from.”
The distinction to carry into the room: telemetry tells you what happened; a gateway can stop it happening. The native controls plus OpenTelemetry give you an inventory and an after-the-fact record, which is most of an audit story and none of an enforcement story.
There is a stable extension for the auth half of this:
Enterprise-Managed Authorization (io.modelcontextprotocol/enterprise-managed-authorization).
The organisation’s IdP becomes the authoritative decision-maker; the client exchanges an identity
assertion for an ID-JAG (Identity Assertion JWT Authorization Grant), then exchanges
that for an access token from the server’s authorization server. Onboarding, revocation and policy
all move to the IdP. It is the right design.
On the official, community-maintained extension support matrix, exactly one client is listed as supporting it: Archestra.AI. Not Claude Desktop, not Claude on the web, not VS Code GitHub Copilot, not Microsoft 365 Copilot, not ChatGPT, not Cursor, not Goose. “Extensions are opt-in and never active by default,” and the docs note that support “typically requires client-level support from the organization’s IT team in addition to the MCP client application.”
So the honest answer to “is the standard solving this?” is: the specification has, and the clients your organisation actually deploys have not. That gap is the entire commercial case for gateways. It is also why the gap will close — which should affect how much you build.
One correction worth making, because the vendor write-ups repeat it: the MCP roadmap (last updated 22 August 2026) does not list gateways or proxies as a priority area. Its five priorities are agentic messaging primitives, HTTP-native transport unification, agent identity and enterprise-ready security, improved primitives, and SDK developer experience. The enterprise work is happening under identity — DPoP, workload identity federation, token exchange — not under a gateway pattern. If someone tells you the protocol is standardising gateways, they have read a blog post rather than the roadmap.
This is the question that separates a decision from a discussion. Figures read from the GitHub API on 14 September 2026.
| Implementation | Licence | Stars | Last push | What it is |
|---|---|---|---|---|
IBM ContextForgeIBM/mcp-context-forge | Apache-2.0 | 4,472 | 14 Sep 2026 | The most substantial open implementation. A gateway, registry and proxy in front of MCP, A2A and REST/gRPC, with federation, plugins and OpenTelemetry tracing |
Docker MCP Gatewaydocker/mcp-gateway | MIT | 1,564 | 26 Aug 2026 | A docker mcp CLI plugin. Containerised server execution with a gateway in front — the closest thing to pinning, because the unit you pin is an image |
Lasso MCP Gatewaylasso-security/mcp-gateway | MIT | 387 | 22 Jan 2026 | Plugin-based orchestration of other MCP servers. Note the date: roughly eight months without a push, in a protocol that revised its specification in that window |
There is also a live commercial category — vendors selling managed MCP gateways with SOC 2 positioning — and cloud API-management products adding MCP support. I have deliberately not ranked those, because at this stage the vendor comparisons are mostly written by the vendors.
Three pieces of primary evidence, all :
And it does not review code: “The MCP Registry delegates security scanning to” the underlying package registries and downstream aggregators. Its namespace verification proves a server came from the domain it claims. That is authenticity of name, not safety of behaviour.
So: buying a gateway is real and the products are real. An internal catalogue of approved servers, the review process that admits things to it, and anything resembling version pinning are still yours to build, whichever product you buy.
Anthropic documents an LLM gateway for
Claude Code: you point ANTHROPIC_BASE_URL at it, it holds the provider credential, issues
one key per developer, and requests arrive at POST /v1/messages. That buys per-developer
attribution, spend control, offboarding and routing. It sits on the inference path and knows
nothing about which MCP servers are connected or what tools were called. An MCP gateway
sits on the tool-call path and knows nothing about your token spend. Two different chokepoints,
one word. Conflating them is the most common way these
conversations go wrong, and noticing the difference out loud is a cheap way to be the most precise
person in the room.
A gateway is an MCP server to the client and an MCP client to the upstream server, which means the authorization security requirements land on it directly. Three normative lines from the 2026-07-28 specification:
Read together, these say something useful: the protocol anticipated the gateway pattern and constrained it before the products arrived. The gateway is not a hole in the standard. It is a component the standard has security requirements for, which is a much better position to buy from — and a conformance question you can put to a vendor.
MCP’s appeal was decentralised, per-developer extension: an engineer adds a server in a minute and gets leverage without asking anyone. A gateway puts a team, a queue and a change-approval process in front of that minute. It also puts a single process in the path of every tool call in the company. When it is down, nobody’s agent can do anything; when it is slow, every workflow is slow; and when it lags the protocol, everyone is on the old protocol.
anthropic-* headers and request bodies verbatim rather than allowlisting”403 on real sessions while short test requests pass” — a central component failing in a way your smoke test cannot see404 until someone updates the gateway’s routing table. Substitute
“new tool” for “new model” and you have the MCP version of the same bottleneck.
I think a gateway is the right answer for most organisations
above a certain size and that it is not a pure win, and anyone selling it to you as a pure win
has not operated one.Same rule as every other mechanism in this course — the cheapest control that does the job, and the job is defined by blast radius, not by tidiness:
| Situation | Cheapest control that does the job |
|---|---|
| Read-only servers over public data | NO GATEWAY A denylist and an inventory. You are buying infrastructure to solve a problem you do not have |
| You need to know what is connected | NO GATEWAY OTEL_LOG_TOOL_DETAILS=1 and aggregate. Inventory is the cheapest audit and it is already there |
| You need a fixed approved set | NO GATEWAY managed-mcp.json, or allowedMcpServers with allowManagedMcpServersOnly: true |
| Servers hold credentials to systems of record | GATEWAY Credential custody is the job, and it is the one no client-side setting can do |
| Tool calls write to production, or touch regulated data | GATEWAY You need enforcement in the path, not a record after the fact |
| Many clients, not just Claude Code | GATEWAY Per-client controls give you N policies and N audit formats. A gateway is the only place one policy can live |
allowManagedMcpServersOnly is user-broadenable; unset ≠ emptyAgent Skills use progressive disclosure across three levels, with a published cost per level:
| Level | When loaded | Documented token cost |
|---|---|---|
| 1 — Metadata | Always, at startup | ~100 tokens per Skill (name + description) |
| 2 — Instructions | When triggered | Under 5K tokens (SKILL.md body) |
| 3 — Resources | As needed | None until accessed |
The promise is explicit: “This lightweight approach means you can install many Skills without context penalty: until a Skill is triggered, only its name and description occupy context.” The architecture is genuinely good, and the filesystem design — scripts run via bash so their code never enters context — is the right idea.
bug and has repro
Opened 20 December 2025, still open, 19 comments. Title: “Skills consume full token count at startup instead of progressive
disclosure (frontmatter only)”. The reporter shows /context attributing
3–5.5K tokens to each skill from the official plugin-dev plugin.
The thread sharpens as more people measure their own installs:
The lesson is not “skills are expensive.” It is that a per-unit figure from a doc is not a budget. Here is the measurement, run against a real 52-skill installation:
Note the two things this shows. First, this installation’s descriptions average ~31 tokens, not the ~100 the docs use — the doc figure is a conservative rule of thumb, and your real floor may be lower. Second, the exposure if progressive disclosure does not hold is 21× the floor, and that is the number worth knowing before an org standardises on a large skill library.
Measure the mechanism in your own installation before you budget for it. Count the
metadata, count the bodies, and know the ratio. Then check /context against the floor. If
they disagree, you have found your number and the doc has not.
This generalises well beyond skills. Any mechanism with a “loads only when needed” story has a best case and a worst case, and the gap between them is the thing to quantify — because that gap is what you are exposed to when the optimisation does not fire.
bug ticket: full descriptions re-injected per turn; ~10K tokens at 148 skillsAnthropic offers two distinct arrangements — zero data retention (ZDR) and HIPAA readiness — and the docs are explicit that they are different things for different purposes: HIPAA readiness “applies a broader set of privacy and security safeguards than ZDR (encryption, access controls, and audit logging) rather than requiring immediate deletion. If your organization handles PHI, HIPAA readiness is the arrangement to use; you do not also need ZDR.”
| If your posture is… | You cannot have… |
|---|---|
| Zero data retention | EXCLUDED Claude Fable 5.1, Mythos 5.1, Fable 5, Mythos 5 — these “require 30-day data retention and are not available under ZDR unless expressly authorized by Anthropic.” Your ZDR posture excludes the most capable models. |
| EXCLUDED Agent Skills — “not covered by ZDR arrangements. Skill definitions and execution data are retained according to Anthropic’s standard data retention policy.” | |
| EXCLUDED Claude Managed Agents — stateful; transcripts persist until deleted | |
| EXCLUDED Claude Console and playground, Claude Teams and Enterprise interfaces, Claude for Excel, consumer plans | |
| EXCLUDED CORS — unsupported for ZDR orgs; browser apps must route through a backend proxy. An architecture constraint, not a policy one | |
| HIPAA readiness (BAA) | EXCLUDED Claude Code — “Claude Code is not covered under HIPAA readiness.” Flat, no configuration |
| EXCLUDED Claude Platform on AWS and Microsoft Foundry | |
| EXCLUDED Beta features generally, unless explicitly listed as eligible | |
EXCLUDED PHI in JSON schemas — structured-output and strict: true schemas are compiled to grammars cached separately and “do not receive the same PHI protections.” Applies to property names, enums, consts and regex patterns |
1. The API enforces this, so it is a runtime failure, not a policy discussion.
A HIPAA-enabled org sending a non-eligible feature gets a
400: “The requested features are not available for HIPAA-regulated organizations
without Zero Data Retention: code_execution.” Your compliance posture will show up as broken
builds.
2. The sharpest exclusion is the model one. An organisation that adopts ZDR as a blanket default has, without necessarily realising it, excluded itself from Anthropic’s most capable models and from Agent Skills. That is a legitimate tradeoff and it should be made deliberately, by someone who knows they are making it — not inherited from a security questionnaire that had a “zero retention” checkbox.
| Artifact | Natural owner | The failure if unowned |
|---|---|---|
| Org-wide context (managed settings, org-tier policy) | Platform / security | Nobody can enforce anything; every control is advisory |
| Repo CLAUDE.md | The repo’s owning team, reviewed like code | Grows additively forever; billed on every turn; no deletion process. This is the default outcome |
| Shared skills | A named reviewer, plus a distribution channel | Duplication, drift, and a skill nobody audited running with your permissions |
| Retrieval index | Data platform | Residency and freshness become nobody’s problem |
| Memory stores | Whoever owns the data classification | A persistent store of model-selected content with no retention policy |
Anthropic’s own guidance is blunt: use Skills only from trusted sources, because “a malicious Skill can direct Claude to invoke tools or execute code in ways that don’t match the Skill’s stated purpose,” and it says to treat installing one like installing software.
So the review is a software review, not a document review, and it has four questions a reviewer must be able to answer:
Two practical constraints on distribution: custom Skills do not sync across surfaces — claude.ai, the API and Claude Code are three separate uploads — and on claude.ai they are per-user with no org-wide admin management. An org standardising on skills is standardising three times, and cannot centrally manage the claude.ai copy at all. Claude Enterprise organisations can turn on Skill content scanning for claude.ai and Cowork uploads, but that scanning does not cover the Skills API or the Console.
The test for whether a practice belongs in your rollout is not whether it sounds responsible. It is whether you can finish this sentence: “if we skip this, here is the specific thing that goes wrong, and here is how we would find out.” A practice that cannot finish that sentence is asking engineers to pay a tax so that a document can exist.
So these are grouped by the axis each one protects, rather than listed. If you only take three, take the first one from each group.
| Practice | Why, and what breaks without it |
|---|---|
| Write down the rate of change before choosing the mechanism | One line in the PR: this content changes <never / quarterly / daily / per-request>. It converts an argument about tools into a question of fact. Without it: the quarterly pricing table lands in a CLAUDE.md, because “quarterly” sounds stable. It is then stale for most of each quarter and billed on every turn of every session that never needed it. This is the single most common mechanism error and it is invisible — nothing fails, it is just wrong and expensive. |
| Put a byte-stability rule on everything above the cache breakpoint | No timestamps, no request IDs, no “Hello, {name}” in the system prompt. Caching is a prefix match; any byte change invalidates everything after it. Without it: a one-line, well-intentioned edit takes your organisation’s cache hit rate to zero with no error and no alert, and the bill rises roughly tenfold on the input side. See Module 2. |
| Re-decide on a cadence, not on complaint | Set a date to revisit each artifact’s placement. Rates of change themselves change. Without it: placement decisions are only ever revisited when someone is annoyed, which selects for the loud problems and never touches the expensive quiet ones. |
| Practice | Why, and what breaks without it |
|---|---|
| Name an owner and a deletion cadence for every context artifact | Ownership is the easy half and most orgs do it. The deletion cadence is the half that is always missing. Without it: a CLAUDE.md grows additively forever, because every addition has an advocate and no removal has one. It is the only mechanism billed unconditionally on every turn, so an unowned one is a permanent, silent, compounding line item. |
| Gate on distribution, not on authoring | Anyone may write a skill or a CLAUDE.md for their own repo with no process at all. Review attaches at the moment other people start depending on it. Without it: you either gate everything — and turn a two-hour improvement into a two-week one, which is the objection in Module 12 and it is correct — or you gate nothing, and an unreviewed artifact ends up running with the whole organisation’s permissions. |
| Prefer the denylist for anything that must never load | “Nothing overrides a denylist match,” while allowlists merge across settings scopes unless allowManagedMcpServersOnly is set.Without it: you have an allowlist in configuration management that a developer can widen in their own settings file, and an enforcement posture that exists only in the document describing it. |
| Practice | Why, and what breaks without it |
|---|---|
Instrument cache_read_input_tokens before you optimise anything else |
It is one field in usage, and it is the difference between paying full input rate and 10% of it on every repeated turn.Without it: you cannot tell a working cache from a broken one, because a broken cache raises no error. Teams have run for months believing caching was on. This is the cheapest instrumentation in the whole stack and the highest-leverage. |
| Budget skill descriptions, not just skill bodies | Descriptions are the always-paid floor; bodies are conditional. On a 52-skill installation the spread between the two was 21× — which means the review question “is this skill worth it?” has two different answers depending on which number you meant. Without it: the unconditional floor grows with every skill anyone adds, nobody attributes the cost to the skills, and the measured-versus-documented dispute in Module 7 decides your bill instead of you. |
| Route latency-tolerant work to Batch, and decide model tier only after caching | Batch is 50% off and the docs state the multipliers stack with other modifiers. Caching and batch are free levers; model choice is a quality tradeoff. Without it: the first cost conversation becomes “should we downgrade the model,” which spends quality to buy savings you could have had for nothing. |
Make your enforcement posture version-aware. Claude Code’s v2.1.259 changed how allowedMcpServers applies to managed servers, and
servers an admin had suppressed that way “start loading on each user’s first launch of
v2.1.259 or later, with no prompt or notice.” On a client
that releases near-daily, a security control you validated once is a security control you validated
against one version. The practice is to test new client releases against your managed configuration
before they reach the fleet — which is the same advice Anthropic gives for LLM gateways, for the
same reason.
Naming these matters more than adding a tenth practice, because each one will be proposed to you and each one sounds obviously correct.
| Commonly proposed | Why it doesn’t earn its place |
|---|---|
| A central prompt library | Prompts are the least stable artifact in the system and the least reusable across codebases. A library of them decays faster than anyone maintains it, and the thing that actually transfers — a working technique demonstrated on your own repository — transfers better as a message in a channel than as a catalogue entry. |
| A per-team token budget | It looks like cost control and behaves like a disincentive to use the tool, which is the opposite of what you were hired to do. Budget the unconditional costs — the prefix, the skill descriptions, the always-loaded context — because those are the ones nobody is choosing. Metered usage that tracks work being done is not the problem. |
| An AI centre of excellence | It centralises the expertise at exactly the moment you need it distributed, and it creates a queue. Anthropic’s own adoption guidance points the other way: the signal that adoption is working is “questions in the channel are being answered by people other than you.” A team whose job is to answer the questions guarantees that never happens. |
| Mandatory training before access | It front-loads the cost and delays the only thing that reliably converts a sceptic, which is one successful result on their own code. Keep the material, drop the gate. |
DORA’s position on AI-assisted development is that “AI’s primary role is as an amplifier, magnifying an organization’s existing strengths and weaknesses,” and that “the greatest returns on AI investment come not from the tools themselves, but from a strategic focus on the underlying organizational system.”
DORA’s follow-up ROI report frames returns as following a J-curve — an initial dip before value — with the dip attributed to the extra effort of verifying AI-generated work. I have read the framing on DORA’s own pages but not the underlying figures, so treat the shape as reported and the magnitude as unquoted.
Two consequences for someone arriving to drive adoption. First, if the surrounding system is weak — slow review, unclear ownership, a flaky test suite — the tool will make that worse before it makes anything better, and you should say so in advance rather than be surprised by it in month two. Second, a leader who expects a straight line up will interpret the normal early dip as failure. Predicting the dip out loud, before it happens, is the cheapest credibility you will ever buy.
The ordering principle is simple: do the steps whose outcome would invalidate later work first. Most rollouts run these in exactly the wrong order, because the visible steps are the fun ones.
| # | Step | Why here, and what it would have cost you later |
|---|---|---|
| 0 | Establish the compliance posture before the pilot, not after it | Module 8’s exclusions are structural, not settings. A ZDR default forecloses the frontier models and Agent Skills; HIPAA readiness excludes Claude Code entirely. Discovering this after a successful pilot is the single most expensive sequencing error available, because you lose the pilot and the credibility that came with it. |
| 1 | One team, real work, a bounded window | Not a bake-off, not a survey. A team whose work you can actually see, doing work that would have happened anyway. The output you need is a decision, and a pilot that cannot produce one is a demonstration. |
| 2 | Instrument the cheap things immediately | cache_read_input_tokens and, if you are on Claude Code, OpenTelemetry with OTEL_LOG_TOOL_DETAILS=1. Both are near-free and both answer questions you will be asked in month three. Retrofitting instrumentation means your first three months have no baseline. |
| 3 | Champions, not a programme | See below. This is where adoption is won, and it is the step most likely to be replaced by a training plan. |
| 4 | Name owners for context artifacts while there are still few of them | Retrofitting ownership onto forty CLAUDE.md files is a different and much worse job than establishing it over four. |
| 5 | Policy and gateways last, sized to actual blast radius | You cannot size a control before you know what people actually connect to — which is what step 2 tells you. Building the gateway first means building it for the usage you imagined. |
Anthropic’s own adoption guidance puts it in one sentence: a concrete, runnable example “removes the gap between curiosity and a first successful use, which is where most adoption efforts stall.” And on the mechanism: “Adoption of a developer tool rarely happens because of a rollout announcement. It happens because someone on the team begins using the tool well, talks about it openly, and makes it easy for others to follow.”
The guidance is specific about what that costs: roughly 15 minutes a week posting wins and prompts, 20 minutes answering questions in a shared channel, 5 minutes running a weekly show-and-tell thread, and 0–30 minutes of optional pairing. That budget is the argument. You are not asking for headcount; you are asking for forty minutes a week from people who are already enthusiastic. A proposal priced that way is very hard to refuse, and a proposal that asks for a programme is very easy to defer.
One organisational point the same guidance makes, which is worth defending: on “what about security and data handling?” it says to refer the question to the administrator, because “champions should not improvise this answer.” That is a correct separation of duties, and it also protects your champions — the fastest way to lose one is to let them answer a compliance question wrongly in public.
Three tests. The first is borrowed and is the best of them.
| Test | Real | Theatre |
|---|---|---|
| Who answers the questions? — this is the vendor’s own week-four success signal: “questions in the channel are being answered by people other than you” | Several people answer; threads resolve without you | Every thread waits for one person, or for a vendor |
| Return versus trial | The same names appear week after week | A large, flat trial count with no repeat cohort — which a “% of engineers who have used it” metric reports as success |
| Does it appear in work that ships? | Usage shows up in merged pull requests and in incident timelines | Usage shows up in a dashboard and a quarterly slide |
The Claude Code analytics dashboard exposes a specific set, and the limits are documented as precisely as the metrics.
| Metric | What it actually says | What it cannot tell you |
|---|---|---|
| Lines of code accepted | Lines the user accepted in-session | Whether any of it survived. The docs state it “excludes rejected suggestions and does not track subsequent deletions” |
| Suggestion accept rate | How often edit-tool suggestions are accepted | Whether accepting was correct. A high rate is equally consistent with good output and with insufficient review |
| PRs with CC / Lines with CC | Merged PRs containing attributed lines | Causation. Attribution matches sessions to PR diffs in a window of 21 days before to 2 days after merge, and code “substantially rewritten by developers, with more than 20% difference, is not attributed” |
| Daily active users / sessions | Engagement | Value. This is the number most likely to be presented as a result |
| Leaderboard | Top contributors by volume | Almost nothing you want. Its documented use is finding people who can help others; the moment it is read as a performance ranking it manufactures exactly the usage it claims to measure |
The docs say the contribution metrics are “deliberately conservative and represent an underestimate… Only lines and PRs where there is high confidence in Claude Code’s involvement are counted.” Say that out loud when you present the numbers. A leader who learns later that a metric was conservative discounts everything you showed them; a leader who is told in advance treats it as a floor.
“Contribution metrics are not available for organizations with Zero Data Retention enabled. The analytics dashboard will show usage metrics only.”
Read that against Module 8. A ZDR posture already costs you the frontier models and Agent Skills. It also costs you the PR-level attribution — which is to say, the compliance posture adopted to get the deployment approved removes the evidence you need to prove the deployment worked. You keep daily active users and lines accepted; you lose the link to shipped work.
This is the same structure as every other exclusion in this course and it is the one most likely to be discovered at the worst moment, because nobody puts “how will we measure this” and “what is our retention posture” in the same meeting. If your organisation is heading for ZDR, agree the alternative evidence — cycle time, review latency, a before-and-after on a named workflow — before it is switched on, while you can still capture a baseline.
Which of these is correct matters more than having a response to each. Three of the six are right, and conceding them quickly is what makes the other three land.
| Argument | Is it right? | The honest reply |
|---|---|---|
| “I’m faster without it.” | OFTEN RIGHT | The vendor’s own guidance concedes this: it is “likely true for code the person writes routinely.” The reply is a redirect, not a rebuttal — suggest the work they avoid: legacy files, unfamiliar services, test scaffolding. |
| “Review will become the bottleneck.” | RIGHT | Generation got cheap and review did not. This is the amplifier finding applied to one queue. If your review capacity is already the constraint, adoption makes it worse first — plan for review, not just for generation. |
| “Governance will slow us down.” | PARTLY RIGHT | Right for personal and team-scoped artifacts, which should have no gate at all. Wrong where blast radius is organisational. Gate on distribution, not authoring (Module 9). |
| “We tried it, it hallucinated.” | USUALLY NOT | Usually a context problem rather than a model problem — the whole of this course is about what was in the window. Re-run the same prompt with the relevant files referenced and the actual error output pasted in. |
| “The window is 1M now, so context management is obsolete.” | NO | Capacity was one of five constraints and the only one that got solved. Filling 1M tokens on Opus 5 is about $5 a turn uncached (Module 3). |
| “Let’s wait for the standards to settle.” | NO, BUT… | It is a reasonable instinct pointed at the wrong layer. Waiting on mechanism choices is sensible — those are vendor-shaped and will be renamed. Waiting on posture, instrumentation and ownership is not, because those are the things that take months and gate everything else. |
Ten situations. For each, pick the cheapest decision moment that still catches the rate of change — not the most powerful mechanism available.
Eight claims about token economics. True or false, on the figures in this course.
Ten things that will be said to you, by people who are not being difficult. For each one, decide whether they are right, partly right, or wrong — before you decide what to say. This is the tool that matters most. Conceding the correct objections quickly is what buys you the standing to push back on the others, and an adoption lead who argues all ten loses the room in the first meeting.
The failures that actually kill deployments are: no owner, no eval set, no definition of done, no rollback, unclear escalation, and a pilot that was never scoped to produce a decision. None of those are context-window problems. An organisation can get every mechanism choice in this course right and still fail, and an organisation can get all of them wrong and succeed because a competent team iterated.
Worse, this material is attractive to the wrong audience: it is tractable, quantitative and feels like progress, which makes it an excellent way to avoid the harder organisational questions. A token-economics review is a more comfortable meeting than “who is accountable when this is wrong.”
Skills, CLAUDE.md, MCP and subagents are not natural categories. They are one vendor’s product decisions, several of which did not exist eighteen months ago and several of which will be renamed or absorbed. Building an organisation’s mental model around them is building on a roadmap. The problem-shaped version is far shorter: you have a budget of tokens per turn; decide what goes in it and who decides. Everything else is this quarter’s packaging.
The entire commercial case for MCP gateways rests on a gap: the standards-track answer exists and the clients don’t implement it. But that gap is the most predictable thing in this course. There is a stable Enterprise-Managed Authorization extension, a Working Group on agent identity, and every major client vendor has an enterprise sales motion that rewards implementing it. An organisation that spends two quarters building an MCP gateway is building a bridge over a river that is being drained, and will then own it — with its on-call rota, its version lag and its outage — for years after the thing it existed to work around has shipped natively.
Look at what is actually available: daily active users, sessions, lines accepted, acceptance rate, attributed lines in merged PRs. Not one of those measures whether the software got better, shipped sooner, or broke less. Even the best of them — PR attribution — is a measure of how much model-produced text survived into a diff, which is closer to a volume metric than a value metric, and it quietly rewards the behaviour it should be suspicious of. Presenting these to a leadership team as evidence of value is the theatre the module claims to be detecting.
Model costs have fallen persistently and engineer time has not. A context optimisation that saves $12,500 a year and costs two engineer-weeks a year to maintain is a loss. Worse, per-turn cost is the metric that is easy to measure, so it attracts attention out of proportion to its size — the right denominator is cost per completed task, which is often dominated by retries, wrong turns and human review.
Module 8 reads as though ZDR-excludes-frontier-models is a law. It is a current product decision that could change with one contract negotiation — the docs even say “unless expressly authorized by Anthropic” and tell you to check your contract terms. An enterprise that hard-codes today’s matrix into its architecture will be wrong in a year.
Review boards for CLAUDE.md files and approval gates for skills are how organisations turn a two-hour improvement into a two-week one. The teams getting value from these tools are the ones iterating fastest, and every control proposed in Module 8 is friction applied to exactly that.
| The objection | Why it fails |
|---|---|
| “Context management is obsolete now that windows are 1M.” | Capacity was one of five constraints and the only one that got solved. Cost, latency, attention and governance all got harder, and cost now scales with adoption. Module 3. |
| “Just cache everything and the cost problem goes away.” | Caching is a prefix match that does nothing for output, and output was 38% of the modelled bill after caching. It is the first lever, not the only one. Module 2. |
| “Put it in the CLAUDE.md” as a default answer. | It is the only mechanism billed on every turn of every session unconditionally, and the one with no natural deletion process. It is the right answer for stable, universal, short content and the wrong one for everything else. Module 5. |
Four things are worth doing on Monday, in order, and none of them is a context-window decision: