Diagrams
The 18 required diagrams, each rendered as ASCII (always readable, great for whiteboards) and Mermaid (renders where supported). Use these to redraw from memory — being able to whiteboard these is a core interview signal.
Index
| # | Diagram | File |
|---|---|---|
| 1 | Transformer inference flow | 01 — Inference Internals |
| 2 | Tokenization to generation | 01 — Inference Internals |
| 3 | Prefill vs decode | 01 — Inference Internals |
| 4 | KV cache | 01 — Inference Internals |
| 5 | PagedAttention | 01 — Inference Internals |
| 6 | Speculative decoding / MTP | 01 — Inference Internals |
| 7 | Local inference stack | 02 — Serving & Gateways |
| 8 | vLLM serving stack | 02 — Serving & Gateways |
| 9 | OpenRouter-style gateway | 02 — Serving & Gateways |
| 10 | LiteLLM-style proxy | 02 — Serving & Gateways |
| 11 | Cursor-style IDE architecture | 03 — RAG, Agents & Coding |
| 12 | VS Code BYOK flow | 03 — RAG, Agents & Coding |
| 13 | RAG architecture | 03 — RAG, Agents & Coding |
| 14 | Agent tool-calling loop | 03 — RAG, Agents & Coding |
| 15 | Model evaluation pipeline | 04 — Eval, Cost, Deploy & Startup |
| 16 | Cost observability pipeline | 04 — Eval, Cost, Deploy & Startup |
| 17 | Production deployment on Kubernetes | 04 — Eval, Cost, Deploy & Startup |
| 18 | Startup product architecture | 04 — Eval, Cost, Deploy & Startup |
How to use these
- Study: trace each box and arrow against the linked phase doc.
- Drill: cover the diagram and redraw it from memory; explain each component out loud.
- Interview: these are the diagrams you'll whiteboard for system-design rounds (interview-prep/).
Mermaid blocks render as diagrams where a Mermaid renderer is available; otherwise they read as labeled flow definitions. The ASCII versions are always whiteboard-ready.