Retention Across 34 Months

Project 1 finishes in month 3. Project 15 needs it in month 34. Nothing in the plan, as originally written, connected those two facts.

The whole budget for retention was five 3-hour rebuild drills — 15 hours against 1,430 hours of acquisition, about 1%. That is not a retention plan, it is a gesture. This page is the actual one, and it costs ~55 hours over the journey (3.8%).


Table of Contents


What Actually Decays

Not everything decays equally, and treating it uniformly wastes the budget. Three classes, in increasing order of durability:

ClassExampleHalf-lifeWorth reviewing?
Arbitrary factsThe efConstruction default; RocksDB's exact fanout; a flag nameWeeksNo. Look it up. That is what documentation is for
ProceduresImplementing beam search; writing a page-table walk; deriving a backward ruleMonthsYes — and only retrieval practice works
ModelsWhy decode is memory-bound; why quorums must intersect; the RUM tradeYears, if built by constructionMostly self-maintaining — but they rot silently into slogans

The failure mode specific to this track is the third row. A model you built by measuring degrades into a phrase you can say. "Decode is memory-bound" survives; the ability to derive \(I = 2b/\text{bytes}\) and explain why batch size is the only variable does not. And the phrase feels identical from the inside — which is the illusion of competence, and it is why self-assessment alone will not catch it.

So the target is procedures and models, not facts. Everything below is aimed at those two.


The Four Mechanisms

#MechanismCostTargetsFrequency
1Review queue~10 min/weekprocedures + modelsweekly
2Rebuild drill3 h × 8proceduresstage boundaries + mid-stage
3Forced reuse0 (already in the plan)procedurescontinuous
4Teaching test~2 h × 6modelsevery ~6 months

Total ≈ 55 hours, versus 15 in the original plan. The marginal 40 hours buy the difference between arriving at P15 with fifteen projects behind you and arriving with three you still remember.


1. The Review Queue

Ten minutes a week. The cheapest and highest-yield of the four.

Every project produces items. An item is a question with a checkable answer, never a fact to re-read:

Q: Derive the arithmetic intensity of transformer decode. Why does model size not appear?
Q: Why must a leader not commit a previous-term entry by counting replicas? Sketch the
   counterexample.
Q: Given L1=128KB, what block size for a 512x512 fp32 matmul, and why is the measured
   optimum smaller?
Q: A Bloom filter at 24 bits/key shows 0 false positives in 200k probes. What can you
   conclude?

Answer out loud or on paper, before checking. Retrieval is the mechanism; recognition is not. If you look first, you have re-read, and re-reading is measured as low-utility.

Intervals

Standard expanding schedule, with one modification:

1 day -> 3 days -> 7 days -> 21 days -> 60 days -> 180 days
wrong at any interval -> reset to 1 day

The modification: an item you answer correctly but slowly or hesitantly moves back one interval rather than forward. Fluency is the point — a derivation you can grind out in ten minutes is not one you can use in a design review.

Running it

Reuse the review.py from your swe-interview-prep track, or a plain markdown file with a due date per item. The tool matters far less than the ten minutes.

Item budget: 4–6 per project, ~80 total. Resist more. A hundred-item queue becomes a chore, gets skipped, and then does nothing. Pick the derivations you would be embarrassed to fumble.


2. The Rebuild Drill

Three hours, no assistant, no notes, no reference to the original. Then diff.

This is the only mechanism that tests procedural knowledge honestly, because writing code from scratch cannot be faked by recognition.

The schedule — eight drills, not five

The original plan had one per stage boundary. That leaves gaps of up to 29 weeks. Eight drills, placed at both stage boundaries and mid-stage:

WeekRebuildTests
W15Multi-head attention forward passP01
W26Beam search + the distance counterP02
W40Bloom filter, including optimal kP04
W54The SSTable reader and its sparse indexP04
W67Raft's election logic, including the vote ruleP05
W83The watermark tracker with idle-partition handlingP07
W99EMA profile + the full metric suiteP08
W117Blocked matmul, and predict the optimal block sizeP14

Scoring, which is the part that matters

OutcomeReading
Working in 3 hYou own it
Not working, but you can find your own bugsYou mostly own it
Cannot startYou never owned it. Schedule a proper re-read, and add three review-queue items

The third outcome is the one the drill exists to detect, and it will happen at least once. It is information, not failure — an ability you believed you had and do not is far better discovered in week 54 than in week 122 when P15 depends on it.

Rules

  • Delete the original from the screen. Different directory, no tab open.
  • Time-box hard at 3 hours. The drill measures fluency, not persistence.
  • Diff afterwards and write three lines: what you forgot, what you did better, what the original does that you now think is wrong.

That third line is the interesting one. Six months of intervening work sometimes makes your old design look wrong — and occasionally it is, which is a genuine result about your own progress.


3. Forced Reuse

Free, because it is already in the dependency graph. The strongest retention mechanism here, and it costs nothing extra.

A component you must use six months later cannot decay quietly — it fails loudly:

ReusedWhereGap
P02's indexP03 (W27), P08 (W84)12 → 69 weeks
P04's engineP05 (W55), P07 (W76)12 → 33 weeks
P05's fault injectorP06 (W68), P07 (W76), P15 (W118)1 → 51 weeks
P01's modelP13 (W21), P08 (W84)13 → 76 weeks
P13's frameworkP14 (W111)57 weeks
EverythingP15 (W118+)up to 110 weeks

Exploit this deliberately. When a later project pulls in an earlier component:

  1. Do not read your old code first. Try to use it from memory of its interface.
  2. Where the interface surprises you, that is a decayed item → review queue.
  3. Where you have to change the old component, ask whether the original design was wrong or whether the new requirement is.

Step 1 converts a routine integration into a free retrieval test, and it costs nothing.


4. The Teaching Test

Two hours, every ~6 months. Six times.

Explain one system, from scratch, to an audience that will ask questions — a colleague, a meetup, a blog post with comments open, or a recorded talk you actually publish.

Why it targets models specifically. Teaching forces you to reconstruct the why, and questions attack exactly the joints where your understanding has thinned into a slogan. You cannot answer "but why does batch size affect intensity and model size not?" from a memorised phrase.

MonthTopicNatural venue
M6What a Transformer costs, by sequence lengthInternal brown-bag
M12The three amplifications, with your own numbersBlog post (W43)
M18Why your Raft had a bug and how the checker found itInternal or meetup
M24When offline metrics fail to predict onlineInternal — your team cares about this one
M30What a syscall and a context switch actually costBlog post
M34The integrated system and its resultConference or the paper's talk

These coincide with the publication timeline in portfolio.md, so five of the six are already scheduled work. The addition is doing them live, with questions, rather than only in writing.

The question you cannot answer is the deliverable. Write it down. It is a review-queue item and possibly a re-read.


The Schedule

Everything above, on one calendar. Additions to the existing plan are marked +.

WhenActivityCost
Weekly, in the review sessionReview queue10 min
End of every projectAdd 4–6 queue items15 min
W15, W40, W54, W83, W99+ Mid-stage rebuild drills3 h each
W26, W67, W117Stage-boundary rebuild drills3 h each
M6, M12, M18, M24, M30, M34+ Teaching test (live)2 h each
Every integrationUse the old component from memory first0
M7, M15, M22, M26, M31Re-take calibration C3 + C530 min

Total: ~55 hours, 3.8% of the journey. For comparison, the writing allocation is 10% and the reading allocation is 15%. Retention at under 4% is not extravagant; it was at 1%.


What Not To Try To Retain

Actively let these go. Trying to hold everything is why retention plans get abandoned.

  • Parameter defaults and flag names. Look them up. efConstruction=200 is not knowledge.
  • API surfaces of libraries. Including your own from two years ago — that is what the README is for.
  • Exact numbers. Retain ratios and derivations. "L1:L2:DRAM ≈ 1:6.5:133" and the ability to re-derive; not "0.91 ns".
  • Paper details beyond the one idea. The extraction note is the artifact; the paper is re-findable.
  • Anything from a project whose exit criteria you cut. If you shipped without leveled compaction, do not carry review items about it.
  • Syntax. Four languages over 34 months means constant lookup, and that is correct.

The test: would a competent engineer look this up rather than recall it? Then let it go. Retention is for the things you must have available while reasoning, not the things you must have available.


Measuring Whether It Works

Retention is invisible without measurement, which is why most retention plans quietly stop.

Three signals, all already collected:

  1. Rebuild-drill outcomes. Working / bugs-findable / cannot-start, per drill. A trajectory toward "cannot start" is the alarm.
  2. Calibration C3 and C5, re-taken at each stage review. Both measure skills the track claims to build. A flat score across two stages means the program is not working for you, and the response is to change the program.
  3. Forced-reuse friction. When P08 pulls in P02's index in week 84, how long before it is working? Note it. Under an hour means the interface was good and you remembered; half a day means something decayed.

Record all three in notebook/retention-log.md. Ten lines a year.

If the drills are all passing and C3/C5 are rising, cut this page's budget — reduce to five drills and four teaching tests and reinvest the hours. Retention effort should be sized to measured decay, not to anxiety about decay.


References

  • Ebbinghaus, H. Memory: A Contribution to Experimental Psychology. 1885. The forgetting curve, and the finding that spacing repetitions beats massing them.
  • Roediger, H. L., Karpicke, J. D. Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention. Psychological Science 17(3), 2006. Retrieval practice beats re-study, by a wide margin, at long delays.
  • Karpicke, J. D., Roediger, H. L. The Critical Importance of Retrieval for Learning. Science 319(5865), 2008.
  • Cepeda, N. J. et al. Distributed Practice in Verbal Recall Tasks: A Review and Quantitative Synthesis. Psychological Bulletin 132(3), 2006. Optimal spacing scales with the retention interval you need — which for a 34-month journey means intervals out to months, not days.
  • Bjork, R. A., Bjork, E. L. Desirable Difficulties in Theory and Practice. JARMAC 9(4), 2020. Why the rebuild drill must be hard and unaided to be worth doing.
  • Koriat, A., Bjork, R. A. Illusions of Competence in Monitoring One's Knowledge. Journal of Experimental Psychology 31(2), 2005. Recognition feels like recall — the empirical reason mechanism 2 exists at all.
  • Dunlosky, J. et al. Improving Students' Learning With Effective Learning Techniques. Psychological Science in the Public Interest 14(1), 2013. Practice testing and distributed practice rank high utility; re-reading and highlighting rank low.
  • Anderson, J. R. Acquisition of Cognitive Skill. Psychological Review 89(4), 1982. The procedural/declarative distinction underlying What Actually Decays.