Stage 5 of the path, and the other one the source list left out. Two phases with opposite bottlenecks, one derivation that tells you your arithmetic intensity is your batch size, and the reason the KV cache rather than the weights usually decides how many users a GPU can serve. Ends with the only cost model you need — and the ranking of levers, which is not the one that gets executive attention.
Stage 5 of five · prerequisite: The Training Stack · the path is in Intermediate to Advanced AI
This stage has the highest ratio of money to understanding of the five. Nearly every serving optimisation is a consequence of one fact — that generating a single token requires reading every parameter from memory — and once you hold it you can usually derive the technique rather than memorise it. Module 2 is that fact, made quantitative.
Every number in the tool is derived from published hardware specifications and the model arithmetic from Stage 1. Where the simulator makes a simplification that would flatter the result, it says so at that point rather than in a footnote.
Two claims in wide circulation are corrected here: that quantization roughly doubles throughput (it halves the bytes, which is not the same thing), and that speculative decoding is a general-purpose speedup (it mostly evaporates at the batch sizes that make serving cheap).
The whole prompt goes through the model at once. Every token is a column in a large matrix multiply, so each weight loaded from memory is used across hundreds or thousands of tokens.
High arithmetic intensity. Compute-bound. The GPU is genuinely busy.
Controls time to first token (TTFT).
One token at a time, each depending on the last. To produce a single token you must read every parameter in the model from memory.
Terrible arithmetic intensity. Memory-bandwidth-bound. The FLOPs sit idle.
Controls time per output token (TPOT).
Sit with the decode case, because it is the one that costs money. A 70-billion-parameter model in BF16 is 141 GB of weights. Every generated token requires reading all 141 GB. At an H100's 3.35 TB/s that is 42 milliseconds of pure memory traffic per token — set entirely by bandwidth, with the arithmetic units almost idle. Double the chip's FLOPs and nothing changes.
In practice you shard those weights across several GPUs, and the read parallelises: on four H100s each device reads its own quarter, so the floor drops to about 10.5 ms. That is a real speedup and it does not change the character of the problem — you are still waiting on memory, just on four memory buses at once.
Nearly every inference optimisation is an attempt to get more useful work out of each byte read from memory.
Batching: read the weights once, produce many tokens. Speculative decoding: read the weights once, produce several tokens. Quantization: make the bytes smaller. Prefix caching: skip reading tokens you already processed. PagedAttention: stop wasting the memory you would have used for more batch.
Five techniques, one idea. Once you have it, you can usually derive the technique from the problem rather than remember it from a paper, and — more useful in practice — you can predict which one will help your deployment before you try it.
A server doing both phases at once is making a compromise it cannot escape. A long prefill occupies the GPU for tens or hundreds of milliseconds, during which every sequence in the decode batch stalls. Users see it as inter-token jitter: fluent output, then a pause, then fluent output.
Two responses are standard, and it is worth knowing both because they represent different philosophies:
Disaggregation is the more interesting of the two, because it is an admission that these are different machines' workloads wearing one model's name. If you internalise nothing else from this module, internalise that: "LLM inference" is two jobs, and every design decision downstream depends on which one you are optimising.
Stage 2 established the ridge point: a kernel is compute-bound above a certain number of FLOPs per byte read, and that threshold is peak FLOP/s divided by memory bandwidth. For an H100 SXM at 989.5 dense BF16 TFLOPS and 3.35 TB/s, the ridge is about 295 FLOPs per byte.
Now compute decode's intensity. One decode step, batch size B, model with N parameters at bp bytes per parameter, ignoring the KV cache for a moment:
That is the whole result, and it is worth stopping on. In BF16, your arithmetic intensity during decode is numerically equal to your batch size. N cancels — a bigger model does not change your position on the roofline at all, it just moves both terms together.
The ridge on an H100 is 295. So you need a batch of roughly 295 concurrent sequences before decode stops being bandwidth-bound. Below that, the arithmetic units are idle in direct proportion to how far below you are.
Batch 1 — a single user, a local model, an agent loop with no concurrency — runs at an intensity of 1 against a ridge of 295. You are using roughly a third of one percent of the chip's arithmetic capability. That is not a bug and not something to fix by writing better kernels: it is the shape of the problem, and the only fix is more sequences in flight.
It also tells you why the local-model experience feels the way it does. Your GPU is not working hard. It is queuing at the memory bus.
The derivation above ignored the KV cache, and at realistic context lengths the KV cache is where the bytes are. Put it back in:
where S is the average sequence length. Notice what happens as B grows: the numerator grows linearly and so does the second term in the denominator. Intensity saturates. Past the batch size at which KV traffic exceeds weight traffic, adding more sequences stops improving your efficiency — you are now reading more cache per step in exact proportion to the extra tokens you produce.
Llama 3 70B in BF16 on four H100s. From Stage 1, its KV cache is 320 KiB per token. Take 4,096-token sequences and the largest batch that fits — Module 3 works this out as about 125 sequences once you leave room for activations.
Two things fall out, and both are worth carrying. First, the KV cache is reading more bytes than the weights are — 168 GB against 141 GB per step. Everyone knows the KV cache occupies memory; far fewer people notice it also consumes bandwidth, which is the scarcer resource.
Second, this deployment is at maximum batch and still more than five times below the ridge. It cannot become compute-bound at this context length, at any batch size, because the thing that would let it — more sequences — brings its own bandwidth cost. Long context does not merely consume memory; it caps your achievable efficiency.
Generating token t+1 requires attention over all previous positions. Recomputing every key and value from scratch at every step would make decoding quadratic in output length, so you cache them. That is what makes generation tractable at all.
From Stage 1, the size is fully determined by the config:
For Llama 3 70B in BF16: 2 × 80 × 8 × 128 × 2 = 327,680 bytes, or 320 KiB, per token. Every token of every sequence in flight.
Four H100s, 320 GiB total. Weights take 131.4 GiB (141 GB in decimal units — the simulator in Module 10 reports binary, so the figures here are stated in GiB to match). Leave 32 GiB for activations and workspace and you have about 157 GiB for KV cache.
| Average sequence length | KV per sequence | Max concurrent sequences |
|---|---|---|
| 1,024 tokens | 0.31 GiB | ~501 |
| 4,096 tokens | 1.25 GiB | ~125 |
| 32,768 tokens | 10.0 GiB | ~15 |
| 131,072 tokens (advertised max) | 40.0 GiB | 3 |
Advertised context length and achievable concurrency are competing for the same memory, and the exchange rate is fixed by a number printed in a public JSON file.
A model advertised at 128K context, served on four H100s, can hold three such sequences at once. Not three hundred. Three. Your per-user cost at full context is therefore roughly a hundred and seventy times your per-user cost at 1K context, for the same model on the same hardware — and that ratio is arithmetic, not an implementation detail you can engineer away.
This is the calculation to run before agreeing to a long-context product commitment. It takes ninety seconds and it is frequently the difference between a feature and a business.
| Technique | Mechanism | Saving | Cost |
|---|---|---|---|
| Grouped-query attention | Query heads share key-value heads. | Exactly the sharing ratio — 8× for Llama 3 70B's 64:8. | A small, generally accepted quality cost. Decided before training; not available to you afterwards. |
| KV cache quantization | Store K and V in FP8 or INT8 instead of BF16. | 2× for 8-bit. | Available at serving time and cheap. Quality impact is usually small but is workload-dependent — measure it on your traffic rather than trusting a benchmark. |
| Eviction / sliding window | Drop or compress cache entries for distant positions. | Unbounded in principle. | You are discarding information the model could have attended to. Sometimes free, sometimes silently destructive on tasks needing long-range recall. Needs its own eval. |
Note that the first is an architecture decision made before training, which is why Stage 1 flagged grouped-query attention as an inference-cost choice baked into the model. The second is the one available to you today at the lowest risk. The third is where the research is and where the unexamined quality regressions live.
Reading 141 GB of weights to produce one token costs you a full memory sweep per token. Reading the same 141 GB to advance 64 sequences costs the same sweep and produces 64 tokens. You have divided the dominant cost by 64, and the extra arithmetic was going to sit idle anyway.
This is why batching is the first lever and the largest, and why a serving system running at low batch occupancy is burning money in a way that no amount of kernel optimisation will recover. It is also why the intensity result in Module 2 matters: batching is not one optimisation among many, it is the movement along the roofline.
Naive batching runs a batch to completion. Requests arrive at different times and finish at different lengths, so a batch of 32 in which 31 sequences finish early runs at 1/32 utilisation until the straggler finishes — and new arrivals queue behind it even though the hardware is nearly idle.
Continuous batching schedules at token granularity. When a sequence emits its stop token, its slot is freed and a queued request takes it on the very next step. Same hardware, several times the throughput, no change to the model.
The idea is Orca (Yu et al., OSDI 2022), which named it iteration-level scheduling: the scheduler invokes the engine to run a single iteration on the batch rather than a whole request to completion, so the batch can be adjusted between iterations. It pairs with selective batching, because not every operation in a transformer can be batched across sequences of different lengths — attention cannot, the rest can.
"Continuous batching" is the name the idea acquired in industry. Knowing the original terms is useful when reading the literature, and the paper is short.
Two real costs, and being clear about them is what separates a serving decision from a slogan.
The naive implementation allocates a contiguous chunk of KV cache per request, sized for the request's maximum possible sequence length — because the cache must be contiguous and you do not know in advance how long the output will be. The PagedAttention paper (Kwon et al., 2023) identifies three distinct wastes in that scheme:
The paper's Figure 2 profiling of existing systems — Orca and FasterTransformer — found that only 20.4% to 38.2% of KV cache memory was being used to store actual token states, and that effective memory "can be as low as 20.4%."
Read that against Module 3. KV cache capacity sets your maximum concurrency, and up to four fifths of it was being thrown away by an allocation strategy. This was not a small optimisation available to a tuned system; it was the dominant inefficiency in production LLM serving, and it was a memory management problem rather than a machine learning one.
Do exactly what an operating system does for process memory. Allocate the KV cache in small fixed-size blocks. Keep a per-sequence block table mapping logical positions to physical blocks. Hand out blocks on demand as the sequence grows.
Three consequences follow immediately:
The reported result is 2–4× throughput at the same latency against FasterTransformer and Orca, with larger gains for longer sequences and larger models — which is what you would predict, since those are the cases where the wasted cache was largest.
Variable-length objects with unknown lifetimes, allocated contiguously in a fixed pool, causing fragmentation. Solved by indirection: fixed-size blocks plus a translation table, which additionally enables sharing.
That is virtual memory, and it was solved in the 1960s. The contribution of PagedAttention is not inventing paging; it is recognising the shape of the problem. This is the clearest recent example in ML systems of a large win coming from knowing what an adjacent field already solved, and the reason to internalise it as a pattern rather than a technique is that the next such win will look different and rhyme.
If two requests share a prefix — the same system prompt, the same few-shot examples, the same retrieved document, the same conversation so far — then the keys and values for that prefix are identical. They depend only on the tokens and their positions, both of which match.
So compute them once and reuse them. Thanks to PagedAttention's block tables, "reuse" means pointing a second sequence's block table at the same physical blocks. Nothing is copied and no extra memory is consumed.
Be precise about the saving, because it is often overstated. Prefix caching eliminates prefill work for the shared portion. It does nothing for decode. If your workload is a 2,000-token system prompt and a 500-token answer, prefill was a meaningful fraction of the cost and you have just removed most of it. If it is a 100-token prompt and a 2,000-token answer, you have saved almost nothing, because decode was always the bill.
Matching is on the token prefix. It is a literal prefix match, from position zero. Anything that varies must come after everything that is shared.
Which means a timestamp, a request ID, a user's name or an A/B flag placed near the top of the system prompt destroys the cache for every request. Nothing errors, nothing is logged, and your prefill cost silently reverts to uncached. This is a genuinely common and expensive mistake, and it is invisible unless you are watching your cache hit rate.
Order your prompt: fixed content first, then slowly-varying content, then per-request content, then user input. That single ordering rule is worth more than most prompt engineering, and it is the sort of thing that never appears in a prompting guide because it is a systems consideration wearing a prompting costume.
A small, cheap draft model proposes the next k tokens. The large target model then verifies all k in a single forward pass — because verification is a prefill-shaped operation: you are scoring a known sequence, not generating one. Accepted prefixes are kept; at the first rejection you resample that position from a corrected distribution and discard the rest.
You have paid one read of the large model's weights and received up to k+1 tokens instead of one. Another bandwidth trade, and the cleanest one in the catalogue.
This is not an approximation and it is worth understanding why. The acceptance test uses a modified rejection-sampling scheme: accept a drafted token with probability min(1, p_target/p_draft), and on rejection sample from a specific residual distribution. The construction guarantees the resulting samples are drawn from exactly the target model's distribution.
Leviathan, Kalman and Matias (ICML 2023) report 2×–3× acceleration on T5-XXL with, in their words, identical outputs. A speedup with no quality cost is rare enough to be worth being suspicious of, and in this case the suspicion does not survive the proof. The draft model's quality affects only how often you accept, never what you produce.
With acceptance rate α and a draft of k tokens, the expected number of tokens accepted per verification pass is the geometric sum
At α = 0.7 and k = 4 that is about 2.8 tokens per verification. Subtract the draft model's cost — typically a few percent of the target's per token — and you have a roughly 2.3× speedup for a single sequence.
The acceptance rate is where the engineering is. It rises with draft-model quality and with how predictable the text is — boilerplate, code with strong conventions and formatting are drafted well; genuinely novel reasoning is not. This is also why self-speculative approaches, n-gram lookup against the prompt, and Medusa-style extra heads all exist: they are different bets about where cheap predictability lives.
Speculative decoding converts idle FLOPs into tokens. Verifying k+1 tokens per sequence costs roughly (k+1) times the arithmetic of a normal decode step, while costing about the same bandwidth. That is a wonderful trade when the arithmetic units are idle.
Module 2 told you exactly when they stop being idle: at a batch size approaching the ridge point. A high-throughput server already running at large batch is much closer to compute-bound, and multiplying its arithmetic by (k+1) makes compute the binding constraint. The speedup shrinks, and at sufficiently large batch it inverts — you are now doing five times the arithmetic for 2.8 times the tokens.
So the honest summary: speculative decoding is a latency technique for low-concurrency serving, not a throughput technique for high-concurrency serving. It is superb for a single user on local hardware, for interactive coding assistants, for latency-sensitive agent loops. It is often close to worthless on a busy multi-tenant endpoint that is already batching well — which is precisely the deployment whose bill you were trying to reduce.
The simulator in Module 10 lets you watch the speedup collapse as you raise the batch size. It is the most counterintuitive result in the tool and the one most worth generating yourself.
Decode is bandwidth-bound and the dominant traffic is the weights. Halve the bytes per parameter and you halve the bytes read per token, so you halve the decode bandwidth floor. Recall the intensity formula: intensity = 2B/bp. Dropping from BF16 to FP8 doubles your arithmetic intensity at the same batch size, moving you along the roofline for free.
Two consequences follow immediately, and the second is the one people get wrong:
You will read that FP8 "roughly doubles throughput." That does not follow from the mechanism and it is usually an overstatement.
Halving the bytes halves the weight-read term. But at any useful batch size the KV cache is also consuming bandwidth (Module 2: 180 GB against 141 GB in the worked example), and quantizing the weights does nothing for it unless you quantize the KV cache too. Meanwhile prefill, scheduling overhead and the sampling step do not shrink at all. End-to-end throughput gains are real and typically well below 2×, with the exact figure depending on your prompt-to-output ratio and your batch size.
Three cases, and conflating them is where the overstatement comes from:
The right way to state it: FP8 halves the bytes you quantize, and separately frees memory that raises your batch ceiling. Whether that reaches 2× depends on whether you quantized the KV cache too and on how decode-dominated your workload is — both of which you can check in the simulator rather than assume.
Quantization quality claims are usually made by whoever is selling the throughput, so the independent work matters. The largest public study is "Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization, which evaluated the whole Llama-3.1 family across more than 500,000 evaluations on academic benchmarks and real-world tasks. Its findings:
| Format | Reported quality | What it buys |
|---|---|---|
| FP8 (W8A8-FP) | "Effectively lossless across all model scales." | Half the weight bytes; native FP8 matmul on H100-class hardware. |
| INT8 (W8A8-INT) | Well-tuned INT8 shows 1–3% accuracy degradation. | Half the weight bytes on hardware without FP8 support. |
| INT4 weights (W4A16) | Competitive with the 8-bit approaches. | Quarter the weight bytes. Activations stay 16-bit, so compute does not speed up. |
The practical reading: FP8 is close to a free lunch on hardware that supports it and should probably be your default. INT4 weight-only is the right choice when memory is the binding constraint — fitting a larger model on fewer GPUs, or freeing memory for KV cache — and note that it does not accelerate compute, so it helps decode and not prefill.
One caveat the study's framing invites and does not remove: benchmark-average quality is not your quality. If your task is unusual, or safety-critical, or has a long tail that benchmarks do not probe, run your own eval — which is Stage 4, and this is one of the cleanest cases for why Stage 4 comes before Stage 5 in the path.
Every serving optimisation moves one of those three terms. Being explicit about which one, and by how much, is what turns a list of techniques into a decision.
The ordering is the point. The largest lever is a scheduling decision. The second-largest is gated on your eval quality. The one that reliably gets a meeting is the smallest by an order of magnitude.
Routing is where inference economics and evals meet, and it is the cleanest justification for why this course is Stage 5 and evals are Stage 4.
To route, you must decide per request whether the cheap model is good enough. That is a classifier, and it has the same two error rates as any other. Route too aggressively and quality falls in ways your users notice and your dashboard does not. Route too conservatively and you have built a complicated system that saves nothing.
The width of that safe operating band is the quality of your eval. A team with a validated judge and a corrected quality number can find the routing threshold empirically and defend it. A team without one is guessing, and will either give the saving back or damage the product. You cannot capture the second-largest lever in inference economics without having done Stage 4.
These figures are order-of-magnitude guidance rather than measurements, and the honest framing matters. The 5–10× for batch efficiency is what you see moving from a naive implementation to a well-configured modern serving stack — and it is consistent with the published numbers behind it: 2–4× from PagedAttention alone against systems that already did continuous batching, on top of what continuous batching bought over naive batching. If you are already running a well-tuned vLLM-class stack, most of that has been collected and your remaining headroom is much smaller.
The 3–5× for routing depends entirely on your traffic mix. If 80% of requests are easy, routing is transformative. If your traffic is uniformly hard, it is worth nothing, and no amount of engineering changes that. Measure your mix before planning around this number — which, again, requires Stage 4.
Commit to a prediction before the simulator runs. That is the clearing test for this stage — being right for the right reason, before you measure — and this is a cheap place to practise it.
Derived, not fitted: parameter counts and KV-cache sizes come from the published configs and reproduce the model cards. Peak FLOPS are the dense datasheet figures, not the sparsity numbers. Decode time is max(bandwidth time, compute time) with both terms computed explicitly, which is what produces the compute-bound transition at high batch rather than assuming it.
Optimistic: it assumes perfect batching with no queueing delay, perfect tensor-parallel scaling with no communication cost, and a fixed MFU that does not degrade at small batch. Real systems are worse than this on every axis. Treat the absolute numbers as a floor and the ratios — which lever wins, by how much — as the useful output.
Not modelled: request arrival variance and the queueing that follows, preemption, chunked prefill interleaving, MoE routing (which changes decode economics substantially, since only the active experts are read but all of them must be resident), and the cost of the draft model's own KV cache under speculative decoding.
Prices: the $3/GPU-hour used for cost figures is nanochat's stated anchor. Rental prices move constantly; scale the cost line accordingly.
Build a toy serving loop with a block-based KV cache and continuous batching — a few hundred lines, no vLLM. Measure tokens per second against a naive implementation that pre-allocates a contiguous cache sized to the maximum length and batches to completion. Report the throughput difference and the memory-utilisation difference, and say which of the two changes produced most of the gain.
Then the diagnosis, which is the real test:
Given a deployment — model size, average prompt length, average completion length, concurrency, hardware — say whether it is prefill-bound or decode-bound, predict which of batching, prefix caching, quantization or speculative decoding will help most and by roughly how much. Then run it and check.
You have cleared Stage 5 when you are right for the right reason. Being right because you tried everything is not the capability. The specific things you should be able to answer unaided:
Not the test: deploying vLLM successfully. vLLM is excellent and it is very good at hiding exactly the mechanics this stage is about — a virtue in production and a problem for learning. Build the bad version first, so the good version means something.