The Mechanism — What a Transformer Actually Computes
Part 1 of a five-part series on how LLMs work. The forward pass step by step: attention, the residual stream, and where the parameters live.
Press ⌘K anywhere to open
Interactive courses I've put together with AI tools, mostly on AI and engineering.
Part 1 of a five-part series on how LLMs work. The forward pass step by step: attention, the residual stream, and where the parameters live.
Part 2. Scaling laws worked out as unit economics: training compute, memory per parameter, parallelism, and why Chinchilla's headline fit didn't replicate.
Part 3. SFT, RLHF and DPO, with DPO's closed form derived step by step, and what that shortcut costs.
Part 4. Read traces and label the failures before picking a metric, then check an LLM judge like any other classifier before trusting its pass rate.
Part 5. Why decoding is limited by memory bandwidth, why the KV cache caps how many users a server can hold, and what that does to serving costs.
Based on Thariq's July 2026 article on how prompting changes for Claude 5 models, with applied modules on CLAUDE.md, skills and tool design.
Skills, CLAUDE.md, memory, retrieval, MCP and subagents compared as different ways to decide what the model sees, and what each one costs per turn.
How Claude Code's extension model is moving from hooks that vote on an action to mods that wrap it, read from a public issue and published source before release.
A close reading of the email exchange Boris Cherny published about AI-written code in an old codebase, and what each side actually argued.
OpenAI's nine-stage agent pipeline, read as a system for deciding what a person should look at rather than a system for writing code.
Anthropic's playbook for running each stage of software delivery with agents, checked against METR, DORA 2025 and other published evidence.
Epoch AI's estimate of how fast the cost of AI is falling, next to Anthropic's proposal for measuring progress from inside a lab.
Dario Amodei's September 2026 essay laid out as a chain of claims, then tested against its sources and the main objections.
TypeSafe AI's Jev returns typed answers with probabilities instead of chat. What its calibration does and doesn't guarantee, and the economic bet behind it.
The engineering in Google's paper on putting TPUs in orbit: power, laser links between satellites, and radiation, with measured results kept apart from projections.