Tracks

Seven tracks, one per capability the loop tests. Each has a concept inventory, a drill set, build artifacts, a failure-mode catalog, and a self-assessment rubric.

No topic is named without a file that teaches it and a drill that tests it. If you find one, that is a bug — log it in ../STATE.md.


Table of Contents


The Seven Tracks

TrackDirectoryTests which roundBaseline share of 570h
A — Coding under time pressurecoding/Technical screen A; onsite Coding 1 & 225%
B — Python internalspython-internals/Onsite Coding 2's follow-ups12%
C — Distributed systems designsystems-design/Technical screen B15%
D — ML & inference infrastructureml-infra/Onsite system design ("design ChatGPT")20%
E — Take-home and deep divethis file + ../projects/The 48-hour build and the line-by-line defense12%
F — Behavioral at staff altitudebehavioral/Recruiter screen; onsite behavioral10%
G — Agentic codingagentic/The beta fifth round6%

Shares are the baseline. They are re-derived from your diagnostic levels per ../diagnostics/RUBRIC.md and rebalanced at every monthly re-test.


How a Track Is Structured

Every track README has the same five sections, so you always know where to look:

  1. Concept inventory — everything the track covers, with the file that teaches each item.
  2. Drill set — what you actually do. Timed, scored, repeatable.
  3. Build artifacts — the things that exist when the track is done. Code, not notes.
  4. Failure modes — how people lose this round, and the specific symptom of each.
  5. Self-assessment rubric — the L0–L3 bands, plus the hire-bar translation.

Track E: Take-Home and Deep Dive

Track E has no directory of its own because its artifacts are real projects. It lives here plus ../projects/.

The reported shape (rows 9–15 of ../research/source-report.md): a 48-hour window to "build something real" — the given example being a distributed webhook delivery system with retry logic and dead-letter queues — followed by a round in which the interviewer walks your code line by line, from a question list he wrote after reading it.

The insight that should reorganize how you build

The take-home and the deep dive are one round, not two. The take-home's real function is to generate a personalized interrogation surface. So:

Every decision you make in the 48 hours is a question you will be asked in week three.

Which inverts the optimization target. It is not "best code." It is "code every line of which I can defend, plus a written record of the alternatives I rejected." A slightly simpler system you can defend completely beats a sophisticated one containing three choices you made on autopilot at hour 31 and cannot now reconstruct.

This is inference I1 in ../research/findings.md — labelled as inference, not sourced. But it follows directly from the reported fact that the interviewer writes the question list after reading your code.

The 48-Hour Playbook

Reported grading criteria converge tightly across sources: code quality, test coverage, a written design doc explaining tradeoffs, and how you handled the deliberately under-specified parts. One source puts it bluntly — a working solution with a thoughtful README beats a clever solution with no docs.

The 48 hours include sleep. Budget them:

HoursPhaseOutput
0–2Read and interrogate the briefA written list of every ambiguity, and the decision you are making about each. This list becomes a README section
2–4Design doc v1Architecture, data model, the two hard parts, what is explicitly out of scope
4–8Walking skeletonEnd-to-end path working with the simplest possible everything. Committed and green
8–28Implementation with tests as you goNot tests at the end. Tests at the end is how you run out of time and ship untested code
28–34Sleep. Non-negotiable
34–40The hard partWhatever you deferred: the failure handling, the concurrency, the benchmark
40–44One thoughtful benchmarkA measured number with the methodology written down
44–47README, design doc v2, commit history cleanup
47–48BufferSomething will be broken. It always is

Never missing, regardless of what you cut:

  • Tests that actually run, with a one-line command to run them
  • A README with run instructions that work on a clean machine
  • A design doc with a tradeoffs section
  • Clean commit history that tells the story of the build
  • Error handling on every external boundary
  • One benchmark with a number and a stated methodology
  • An explicit "what I would do with two more days" section

"Beyond the ask" means — and this is a narrow definition, deliberately: not more features. It means one of (a) a measured benchmark with an honest methodology, (b) a failure-injection test that proves a recovery path actually works, (c) an operational concern nobody asked for but every reviewer notices — structured logs, a health endpoint, a runbook for the DLQ. Anything else is scope creep and it reads as poor judgement.

The Decision Log

Start it at hour zero. Append as you go. It is the single highest-leverage artifact in the whole track, and it costs about ninety seconds per entry.

## D-007 — Retry backoff: full jitter
- **Decision:** exponential backoff with full jitter, base 200ms, cap 30s, 6 attempts
- **Alternatives:** no jitter (rejected: synchronized retry storms after a
  downstream recovery — this is the actual failure mode AWS documented);
  equal jitter (rejected: marginal benefit over full at our concurrency);
  decorrelated (rejected: harder to reason about a worst-case bound)
- **Assumes:** downstream recovery is correlated across our consumers
- **Would revisit if:** we ever have a single-tenant destination where
  ordering matters more than throughput
- **Not tested:** behavior when the clock jumps backwards

At the deep dive you will be asked "why 200ms?" and "why six attempts?" Ninety seconds at hour 12 buys you a complete answer at week 3. Without the log, you will reconstruct a rationalization, and the interviewer will hear it as one.

The Deep-Dive Interrogation Harness

After each project ships, I read your actual diff and generate the question list an interviewer would write. Project-specific, not generic — that is the whole point of row 15.

Question classes, all of which will be asked:

ClassExamples
ChoiceWhy this data structure? Why this library and not the stdlib? Why this concurrency model?
Magic numbersWhy 30 seconds? Why 6 retries? Why a batch of 100? Where did that come from?
ScaleWhat happens at 100x? Which component fails first? What is the first thing that pages?
Data lossWhere can this lose a message? Which crash points are unsafe? What is your durability boundary?
OmissionWhat did you not test? What is the least-tested path? What is the riskiest line in the diff?
RegretWhat would you do with two more days? What would you rip out?
HostileThis function does four things. Why? · This test asserts nothing meaningful. · You catch a bare Exception here. · This is O(n²) and you know it

Then it runs as a live drill: 45 minutes, no notes, recorded, scored on the same hire-bar scale as everything else.

Run it twice, on two different projects. Row 15 is about generalization: if you only ever defend the webhook system, you have memorized answers rather than built the skill.

Track E Rubric

LevelStandard
L0Ships something working; no design doc; tests written at the end or not at all
L1Ships with tests and a README; cannot defend specific constants under questioning
L2Ships with tests, design doc, decision log; defends most choices; some "I'd have to look"
L3Defends every line including omissions; names the alternatives rejected and what would flip each; volunteers the weakest part of the design before being asked

L3's tell: volunteering your own design's weakest point before the interviewer finds it. It is the single most credibility-generating move available in this round, and almost nobody does it, because it feels like arguing against yourself. It is the opposite — it demonstrates you have a model of your own system's risk, which is exactly what they are testing.


Completion Rules

These apply to every track, without exception:

  • Reading something never completes anything. Completion requires a passed drill, a working artifact, or a scored mock.
  • Anything you got wrong enters ../review/ at the 1-day interval and resurfaces at 1, 3, 7, and 21 days.
  • Every performance claim in these notes has a script that demonstrates it. If you find one that does not, it is a bug.
  • Every week ends with a scored mock and a ../STATE.md update.

References