Primary-Source Readings

Every reading in the journey, by project, with the reason and the hour budget.

Total: ~183 hours across 130 weeks — about 13% of the 1,430-hour budget, close to the 15% allocation. The remainder of that allocation goes to reading source code, which is not listed here and is often more valuable than the papers.


Table of Contents


Three Rules

1. Read after designing, not before. Every project's reading is scheduled after the milestone where you write your own version. Reading Malkov before designing your own graph index destroys the reconstruction exercise permanently, and it is the exercise the whole journey is built around.

2. Read the section, not the paper. Most entries below name a section. Raft §5.4.2, Vaswani §3, Dataflow's model section. Reading a paper end to end is usually a worse use of ninety minutes than reading its key section twice.

3. Reading never completes anything. It is an input to a milestone, never a deliverable. A week whose output is "read the LSM paper" is a failed week.


The Method Canon

Six pieces about how to work, not about any system. Read the first two in month 1; the rest as noted.

ReadingWhenHours
Hamming, R. W. You and Your Research. Bell Labs, 1986Week 11
Feynman, R. P. Cargo Cult Science. Caltech, 1974Week 10.5
Lampson, B. W. Hints for Computer System Design. SOSP 1983Month 21.5
Dean, J., Barroso, L. A. The Tail at Scale. CACM 56(2), 2013Before P051
Ousterhout, J. Always Measure One Level Deeper. CACM 61(7), 2018Before P050.5
Mytkowicz, T. et al. Producing Wrong Data Without Doing Anything Obviously Wrong! ASPLOS 2009Before your first speedup claim1

Mytkowicz is the one people skip. It shows that changing the link order of a program, or the size of an environment variable, shifts measured performance by enough to invert published conclusions. Read it before you trust your first 5% improvement.


P01 — Transformer

13 hours. Full context: P01.

ReadingWhyWhenh
Vaswani, A. et al. Attention Is All You Need. NeurIPS 2017§3 closely. Note it is post-norm — later reversedWeek 2, after your naive design3
Su, J. et al. RoFormer. arXiv:2104.09864, 2021§3.2 derives the relative-position propertyWeek 62
Xiong, R. et al. On Layer Normalization in the Transformer Architecture. ICML 2020Why pre-norm won, via gradient magnitude at initWeek 7, before E62
Dao, T. et al. FlashAttention. NeurIPS 2022§2–3 only. IO-awareness — the same lesson as P14Week 72
Loshchilov, I., Hutter, F. Decoupled Weight Decay Regularization. ICLR 2019Why L2 ≠ weight decay under AdamWeek 51.5
Radford, A. et al. Language Models are Unsupervised Multitask Learners. 2019The GPT-2 architecture tableWeek 41
Sennrich, R. et al. Neural Machine Translation of Rare Words with Subword Units. ACL 2016BPEWeek 81
Ba, J. L. et al. Layer Normalization. arXiv:1607.06450, 2016SkimWeek 40.5

Also worth reading, after milestone 7: Elhage, N. et al. A Mathematical Framework for Transformer Circuits, Anthropic 2021 — the residual-stream-as-a-bus view, which changes how you read every subsequent architecture.


P02 — ANN Index

11 hours. P02.

ReadingWhyWhenh
Malkov & Yashunin. HNSW. IEEE TPAMI 42(4), 2020The source. Algorithm 4 is the part that matters most and is skipped mostWeek 12, after your NSW3
Malkov et al. NSW. Information Systems 45, 2014The single-layer version you will have reinventedWeek 111.5
He, Kumar & Chang. On the Difficulty of Nearest Neighbor Search. ICML 2012Relative contrastWeek 9, before generating data1.5
Beyer, K. et al. When Is "Nearest Neighbor" Meaningful? ICDT 1999The concentration result underneathWeek 91.5
Aumüller et al. ANN-Benchmarks. Information Systems 87, 2020The evaluation protocol to imitate exactlyWeek 101.5
Jégou, Douze & Schmid. Product Quantization. IEEE TPAMI 33(1), 2011The other major familyWeek 142

P03 — Vector Database

11 hours. P03.

ReadingWhyh
Subramanya, S. J. et al. DiskANN. NeurIPS 2019The disk-resident answer; read before designing your segments2
Mohan, C. et al. ARIES. ACM TODS 17(1), 1992§1–3 only. WAL, LSN, redo/undo2.5
Crotty, Leis & Pavlo. Are You Sure You Want to Use MMAP…? CIDR 2022Read after your mmap milestone, then re-examine honestly1.5
Gollapudi, S. et al. Filtered-DiskANN. WWW 2023After your own filtering experiment1.5
Wang, J. et al. Milvus. SIGMOD 2021A real system's segment architecture1.5
Pillai, T. S. et al. All File Systems Are Not Created Equal. OSDI 2014What your fsync discipline actually guarantees1
Kleppmann, M. DDIA, ch. 3The clearest storage-engine overview1

P04 — LSM Engine

15 hours — the largest budget, because this literature is unusually good. P04.

ReadingWhyh
O'Neil, P. et al. The Log-Structured Merge-Tree. Acta Informatica 33, 1996The origin; §3's cost model is the amplification derivation3
Dayan, Athanassoulis & Idreos. Monkey. SIGMOD 2017Bloom bits should not be uniform across levels. Genuinely surprising, and your best hypothesis source here2
Rosenblum & Ousterhout. Log-Structured File System. SOSP 1991The cleaning-cost analysis is the compaction analysis2
Dong, S. et al. Optimizing Space Amplification in RocksDB. CIDR 2017Production numbers for your trade2
Ghemawat & Dean. LevelDB source and implementation notesRead after milestone 52
Athanassoulis, M. et al. The RUM Conjecture. EDBT 2016The framing that makes the project one idea1.5
Chang, F. et al. Bigtable. OSDI 2006SSTables in context1.5
Bloom, B. H. Space/time trade-offs in hash coding… CACM 13(7), 1970Three pages. Read the original0.5
Pillai, T. S. et al. All File Systems Are Not Created Equal. OSDI 2014Re-read0.5

P05 — Distributed KV

20 hours — the largest in the journey. P05.

ReadingWhyh
Ongaro & Ousterhout. In Search of an Understandable Consensus Algorithm. ATC 2014The extended version. §5.4.2 twice5
Lamport, L. Time, Clocks, and the Ordering of Events. CACM 21(7), 1978Happens-before2
Fischer, Lynch & Paterson. Impossibility of Distributed Consensus… JACM 32(2), 1985Theorem and intuition; proof optional2
Herlihy & Wing. Linearizability. ACM TOPLAS 12(3), 1990The definition your checker implements2
DeCandia, G. et al. Dynamo. SOSP 2007The AP design point2.5
Corbett, J. C. et al. Spanner. OSDI 2012What a bounded clock buys2
Gilbert & Lynch. Brewer's Conjecture… SIGACT News 33(2), 2002CAP as a theorem1.5
Hayashibara, N. et al. The φ Accrual Failure Detector. SRDS 2004For E41.5
Kingsbury, K. Jepsen analyses — pick three real systemsWhat violations look like in shipped software1.5

P06 — MapReduce

11 hours. P06.

ReadingWhyh
Dean & Ghemawat. MapReduce. OSDI 2004§3.6 on backup tasks is what E4 tests2.5
Ghemawat, Gobioff & Leung. The Google File System. SOSP 2003The storage assumptions underneath2
Zaharia, M. et al. Resilient Distributed Datasets. NSDI 2012Why lineage beats re-execution2
Zaharia, M. et al. Improving MapReduce Performance in Heterogeneous Environments. OSDI 2008LATE — naive speculation actively harms1.5
Dean & Barroso. The Tail at Scale. CACM 56(2), 2013Re-read with a straggler in front of you1.5
Isard, M. et al. Dryad. EuroSys 2007The general-DAG generalisation1
Verma, A. et al. Borg. EuroSys 2015Where tasks actually run0.5

P07 — Streaming

12 hours. P07.

ReadingWhyh
Akidau, T. et al. The Dataflow Model. VLDB 2015The most important paper in this project. What/where/when/how3
Akidau, T. Streaming 101 / 102. O'Reilly, 2015The clearest explanation of watermarks in print2
Carbone, P. et al. Lightweight Asynchronous Snapshots. arXiv:1506.08603, 2015Flink's barrier snapshotting2
Chandy & Lamport. Distributed Snapshots. ACM TOCS 3(1), 1985The algorithm underneath it1.5
Zaharia, M. et al. Discretized Streams. SOSP 2013The micro-batch alternative, honestly1.5
Kreps, J. The Log. LinkedIn, 2013Changes how you see storage generally1
Kreps, Narkhede & Rao. Kafka. NetDB 2011The log as a primitive1

P08 — Recommender

10 hours. P08.

ReadingWhyh
Chaney, Stewart & Engelhardt. How Algorithmic Confounding… RecSys 2018The feedback loop, simulated. Sets up P092
Covington, Adams & Sargin. Deep Neural Networks for YouTube Recommendations. RecSys 2016Two-stage architecture; "example age" is the freshness lesson1.5
Steck, H. Calibrated Recommendations. RecSys 2018Why accuracy-optimal recommendations are miscalibrated1.5
Cañamares & Castells. Should I Follow the Crowd? SIGIR 2018Why popularity baselines are so hard to beat1.5
Hu, Koren & Volinsky. Collaborative Filtering for Implicit Feedback Datasets. ICDM 2008The implicit-feedback formulation1.5
Wu, F. et al. MIND. ACL 2020News-specific evaluation and its pitfalls1.5
Carbonell & Goldstein. The Use of MMR… SIGIR 1998Four pages0.5

P09 — Simulator

9 hours. P09.

ReadingWhyh
Chuklin, Markov & de Rijke. Click Models for Web Search. 2015Chapters 3–4. The definitive treatment2.5
Chaney et al. How Algorithmic Confounding… RecSys 2018Re-read; closest published work to this project2
Ie, E. et al. RecSim. arXiv:1909.04847, 2019Read the design decisions, then make your own1.5
Craswell, N. et al. An Experimental Comparison of Click Position-Bias Models. WSDM 2008Where the position-bias exponent comes from1
Rohde, D. et al. RecoGym. arXiv:1808.00720, 2018A second design to compare against1
Jeunen, O. Revisiting Offline Evaluation… RecSys 2019Why offline evaluation fails1

P10 — A/B Testing

9 hours. P10.

ReadingWhyh
Kohavi, Tang & Xu. Trustworthy Online Controlled Experiments. Cambridge, 2020Ch. 1–3, 17–19. The book on this3
Kohavi, R. et al. Online Controlled Experiments at Large Scale. KDD 2013SRM, Twyman's law, the real failure modes1.5
Deng, A. et al. Improving the Sensitivity… (CUPED). WSDM 2013Variance reduction1.5
Johari, R. et al. Peeking at A/B Tests. KDD 2017The principled fix for peeking1.5
Kohavi & Longbotham. Unexpected Results in Online Controlled Experiments. 2010Case studies where intuition lost1
Gupta, S. et al. Top Challenges… SIGKDD Explorations 21(1), 2019What the industry finds hard0.5

P11 — Language

16 hours across both phases. P11.

ReadingPhaseWhyh
Nystrom, R. Crafting Interpreters, Part IIITree-walk. After your milestone 54
Nystrom, R. Crafting Interpreters, Part IIIIIBytecode, VM, GC. After milestone 105
Jones, Hosking & Moss. The Garbage Collection Handbook, 2nd ed.IICh. 2–3 and 93
Wilson, P. R. Uniprocessor Garbage Collection Techniques. IWMM 1992IIThe best survey1.5
Ertl & Gregg. The Structure and Performance of Efficient Interpreters. JILP 5, 2003IIDispatch techniques, measured1.5
Pratt, V. Top Down Operator Precedence. POPL 1973INine pages1

P12 — Kernel

17 hours. P12.

ReadingWhyh
Arpaci-Dusseau & Arpaci-Dusseau. Operating Systems: Three Easy PiecesVirtualization + concurrency. Read alongside your current milestone6
Cox, Kaashoek & Morris. xv6 (RISC-V edition)The book and the source. ~9,000 readable lines4
Lampson, B. W. Hints for Computer System Design. SOSP 1983Re-read; written by an OS designer about OS design1.5
Denning, P. J. The Working Set Model for Program Behavior. CACM 11(5), 1968Why locality makes any of this work1
Ousterhout, J. Why Aren't Operating Systems Getting Faster…? USENIX 1990Still true1
Ritchie & Thompson. The UNIX Time-Sharing System. CACM 17(7), 1974Design taste in eleven pages1
Anderson, T. E. et al. Scheduler Activations. SOSP 1991The user/kernel threading boundary1
Bélády, Nelson & Shedler. An anomaly in space-time characteristics… CACM 12(6), 1969The anomaly you will reproduce0.5
RISC-V Privileged Architecture SpecificationTrap and paging chapters1

P13 — Tensor Framework

12 hours. P13.

ReadingWhyh
Baydin, A. G. et al. Automatic Differentiation in ML: a Survey. JMLR 18, 2018The clearest treatment of modes and their costs3
Paszke, A. et al. PyTorch. NeurIPS 2019Design decisions of the thing you are reimplementing2
Abadi, M. et al. TensorFlow. OSDI 2016The static-graph alternative and its rationale2
Griewank & Walther. Evaluating Derivatives, 2nd ed.Ch. 3–4, for rigour2
Chen, T. et al. Training Deep Nets with Sublinear Memory Cost. 2016Gradient checkpointing1.5
Chen, T. et al. TVM. OSDI 2018Fusion as a compiler problem1.5

P14 — Hardware-Aware

13 hours. P14.

ReadingWhyh
Jouppi, N. P. et al. In-Datacenter Performance Analysis of a TPU. ISCA 2017Read after designing your own accelerator3
Goto & van de Geijn. Anatomy of High-Performance Matrix Multiplication. ACM TOMS 34(3), 2008Why BLAS is fast2.5
Chen, Emer & Sze. Eyeriss. ISCA 2016Dataflow taxonomy; the energy argument2
Drepper, U. What Every Programmer Should Know About Memory. 2007Long; the best treatment of cache behaviour2
Williams, Waterman & Patterson. Roofline. CACM 52(4), 2009The model1.5
Micikevicius, P. et al. Mixed Precision Training. ICLR 2018Why fp16 needs loss scaling1
Dettmers, T. et al. LLM.int8(). NeurIPS 2022Where naive int8 breaks1

P15 — Integration and Writing

~10 hours, plus question-specific literature. P15, final-system.

ReadingWhyh
Blackburn, S. M. et al. The Truth, The Whole Truth, and Nothing But the Truth. ACM TOPLAS 38(4), 2016The best checklist for a systems evaluation section2
Hoefler & Belli. Scientific Benchmarking of Parallel Computing Systems. SC 2015Twelve rules; apply all twelve1.5
Peyton Jones, S. How to Write a Great Research Paper. MSR 2004Write the paper first1
Zobel, J. Writing for Computer Science, 3rd ed.The report format2
Collberg & Proebsting. Repeatability in Computer Systems Research. CACM 59(3), 2016Before the reproducibility appendix1
Bailey, D. H. Twelve Ways to Fool the Masses… 1991A list of things not to do0.5
Shewchuk, J. R. Three Sins of Authors in Computer Science and Math. 19970.5 h, permanently useful0.5
Question-specific literatureVaries1.5+

Source Code Worth Reading

Often more valuable per hour than the papers, and not counted in the budget above. Always read after your own implementation, never before.

CodebaseAfterWhy
micrograd (Karpathy, ~150 lines)P13 milestone 1You will have independently invented most of it
nanoGPT (Karpathy)P01 milestone 7A check on your choices, not a source for them
LevelDBdb_impl.cc, version_set.ccP04 milestone 9The clearest small LSM in existence
xv6 (~9,000 lines)P12, continuouslySmall enough to hold entirely in your head
hnswlibhnswalg.hP02 milestone 7~1,000 lines; the neighbour heuristic in practice
etcd/raftP05 milestone 8A production Raft with a readable state machine
Redis — t_string.c, ae.cany timeExemplary C
SQLite — the source and the documentationP03/P04Possibly the best-documented codebase in existence
CPython — ceval.cP11 phase IIA real dispatch loop
Lua 5.x (~20,000 lines)P11 phase IIA register VM; small, complete, elegant

The Ten That Matter Most

If a quarter goes badly and you must triage, protect these.

#ReadingWhy it survives the cut
1Ongaro & Ousterhout, Raft (extended)The only way to get P05 right, and §5.4.2 is a bug you will otherwise ship
2Akidau et al., The Dataflow ModelReframes "correctness" for unbounded input; nothing else does this
3Jouppi et al., TPUMakes the entire hardware/software boundary legible
4O'Neil et al., LSM-TreeThe cost model that generalises to every storage decision
5Baydin et al., AD SurveyThe forward-vs-reverse argument, cleanly
6Kohavi, Tang & Xu, Trustworthy ExperimentsThe only book here that will change what you do at work next week
7Malkov & Yashunin, HNSWAlgorithm 4, which everyone skips and which is load-bearing
8Dean & Barroso, The Tail at ScaleSix pages; permanently changes how you read a latency number
9Lampson, Hints for Computer System DesignThe closest thing to transferable design judgement in print
10Hamming, You and Your ResearchAbout whether you finish, which is the binding constraint

Nine of the ten are freely available. The exception is Kohavi et al., which is worth buying.