Diagrams

The 18 required diagrams, each rendered as ASCII (always readable, great for whiteboards) and Mermaid (renders where supported). Use these to redraw from memory — being able to whiteboard these is a core interview signal.

Index

#DiagramFile
1Transformer inference flow01 — Inference Internals
2Tokenization to generation01 — Inference Internals
3Prefill vs decode01 — Inference Internals
4KV cache01 — Inference Internals
5PagedAttention01 — Inference Internals
6Speculative decoding / MTP01 — Inference Internals
7Local inference stack02 — Serving & Gateways
8vLLM serving stack02 — Serving & Gateways
9OpenRouter-style gateway02 — Serving & Gateways
10LiteLLM-style proxy02 — Serving & Gateways
11Cursor-style IDE architecture03 — RAG, Agents & Coding
12VS Code BYOK flow03 — RAG, Agents & Coding
13RAG architecture03 — RAG, Agents & Coding
14Agent tool-calling loop03 — RAG, Agents & Coding
15Model evaluation pipeline04 — Eval, Cost, Deploy & Startup
16Cost observability pipeline04 — Eval, Cost, Deploy & Startup
17Production deployment on Kubernetes04 — Eval, Cost, Deploy & Startup
18Startup product architecture04 — Eval, Cost, Deploy & Startup

How to use these

  • Study: trace each box and arrow against the linked phase doc.
  • Drill: cover the diagram and redraw it from memory; explain each component out loud.
  • Interview: these are the diagrams you'll whiteboard for system-design rounds (interview-prep/).

Mermaid blocks render as diagrams where a Mermaid renderer is available; otherwise they read as labeled flow definitions. The ASCII versions are always whiteboard-ready.