Track F — Behavioral at Staff Altitude
Reported: technical leadership, architecture decisions spanning teams, driving consensus under pressure, concrete tradeoffs rather than soft-skills answers (
../../research/source-report.mdrows 33–36), plus a recruiter-screen question about where AI is headed (row 4).And the finding that should change how much time you give this: reported sources name the values/culture round as the leading failure point at a peer lab. For companies with technical bars this high, that is a remarkable claim — and senior engineers systematically under-prepare here because it does not feel like real work.
→ Study guide: WARMUP.md — all twelve story categories with worked model answers at staff density, the probe playbook, and the forward-looking answers written out in full.
Table of Contents
- Why This Is Graded at Staff, Not Senior
- DTAO, Not STAR
- The Story Bank
- Required Story Categories
- Probe Lists
- The Forward-Looking Questions
- Drill Set
- Failure Modes
- Self-Assessment Rubric
- References
Why This Is Graded at Staff, Not Senior
AI-lab levelling is compressed. Reported consensus: an OpenAI level maps roughly one level
higher than the same number at Google or Meta, and L5 — titled "Senior" — carries
Staff-equivalent scope (../../research/findings.md).
The practical consequence for every story you tell:
| Senior story | Staff story |
|---|---|
| "I designed and built X" | "I decided X over Y, and got three teams to agree" |
| "It improved latency 30%" | "It improved latency 30% and cost the storage team 60% more memory, which is why they pushed back" |
| "I convinced them" | "I built a prototype that made the argument for me" |
| "It went well" | "It went well; here is the part I got wrong and what it taught me" |
| One team | Cross-team, with a named opponent |
The reliable test: if nobody in the story disagreed with you, it is not a Staff story. Consequential decisions attract disagreement. A story with no opposition is either not consequential or you have edited the opposition out — and interviewers probe for exactly this.
DTAO, Not STAR
STAR front-loads Situation — the least interesting part — and buries the decision. It was designed for interviews that wanted to know whether you were a good teammate. This round wants to know how you think.
| Letter | Section | Share |
|---|---|---|
| D — Decision | What you decided, in one sentence, first | 1 sentence |
| T — Tradeoff | The alternatives and why each lost. Numbers live here | ~40% |
| A — Alignment | Who disagreed, and what you actually did about it | ~30% |
| O — Outcome | Measured, including what you got wrong | ~25% |
Context goes in a clause, not a paragraph: "On the multilingual ranking pipeline, we decided X." That is enough. If they need more they will ask — and them asking is a good sign, because it means they are engaged rather than waiting for you to finish.
The Story Bank
Twelve to fifteen stories from your actual history. I will not invent them, embellish them, or let a Senior-scope story be presented as Staff-scope.
The raw material I need from you, per story — bullets are fine, prose is not required:
- What was the decision? (one sentence)
- What constraint made it hard?
- What alternatives did you seriously consider?
- Who disagreed, and what did they want instead?
- What did you do to get alignment?
- What was the measured outcome?
- What did you get wrong?
Your history — multilingual search and recommendation, media streaming, networking, enterprise infra, cloud — is unusually rich for this. Ranking pipeline redesigns, index-serving migrations, embedding infrastructure decisions, and cross-org platform migrations are all naturally Staff-altitude if you write the tradeoff rather than the tour.
Each story lands in stories/NN-slug.md with: the DTAO write-up, a 90-second spoken version, a
2-minute version, the tags it covers, and its probe list.
Required Story Categories
Every one must be filled. A gap here is a gap the interviewer will find.
| # | Category | Why it is asked | Status |
|---|---|---|---|
| 1 | Architecture decision affecting multiple teams | Row 34 — the core Staff signal | ☐ |
| 2 | A disagreement you lost | The highest-signal prompt that exists. See below | ☐ |
| 3 | A disagreement you won, and why they conceded | Tests whether you persuade or just outlast | ☐ |
| 4 | An outage you owned | Ownership under pressure; blameless analysis | ☐ |
| 5 | A project you killed or descoped | Sunk-cost resistance. Rare and valuable | ☐ |
| 6 | Mentoring / raising a team's bar | Scope beyond your own output | ☐ |
| 7 | A bet that failed | Calibrated risk-taking, honestly reported | ☐ |
| 8 | Driving consensus without authority | Row 35 | ☐ |
| 9 | Shipping under a hard deadline with quality tension | The tradeoff nobody escapes | ☐ |
| 10 | A time you changed your mind from data | Updates on evidence | ☐ |
| 11 | Working with non-engineering partners | Reported: collaboration with researchers, PMs, safety | ☐ |
| 12 | Something you built that you would now build differently | Technical judgement over time | ☐ |
Category 2 deserves its own note
Four ways it fails, all of them visible from across the room:
| Failure | What it sounds like | What it signals |
|---|---|---|
| Humble-brag | "I lost, but six months later they did it my way" | You cannot actually update |
| Victim | "Management overruled me for political reasons" | You do not distinguish wrong from outvoted |
| Trivial | A disagreement about naming | You avoid consequential conflict |
| Revisionist | "In hindsight they were right about everything" | Performed humility; no real position |
What works: state your position as strongly as you actually held it, state theirs fairly enough that they would recognize it, say what decided it and whether the process was sound even if the outcome was not, and say how you behaved after losing — committed or sandbagged.
Then give your honest current read. All three of these are strong:
- "They were right, and here is what I had not weighted properly."
- "I still think I was right, and here is the evidence that has accumulated since."
- "We were both solving the wrong problem."
Only performed humility is weak.
Probe Lists
Every story needs three follow-ups written out — the questions an interviewer asks to test whether the story is real. Generic examples; each story gets its own specific set:
- "What was their strongest argument?" — the single most discriminating probe in the round. If you cannot produce a strong version of the opposing case, you never engaged with it, and the whole story becomes suspect.
- "What would have had to be true for the other option to win?" — tests whether you modelled the decision or pattern-matched it.
- "Who else was affected that you did not mention?" — tests scope honesty.
- "What did that cost the other team?" — cross-team decisions always cost someone.
- "How long did it take, and how much of that was the disagreement?" — tests whether the consensus story is real.
- "What did you measure, and how did you know it was not a coincidence?" — tests rigour.
- "What would you do differently?" — the answer must be specific, not "communicate more."
The Forward-Looking Questions
Written answers, rehearsed weekly, kept in forward/. Company-specific material lives in
../../research/company-brief.md.
| Question | Length | The bar |
|---|---|---|
| Where is AI headed? | 90s | A specific falsifiable claim + evidence + a falsifier + what you would build |
| Why this company? | 30s | Grounded in the work, not the brand |
| What would you work on? | 30s | Concrete, and connected to what you have shipped |
| What is their hardest unsolved engineering problem? | 60s | A real technical position you can defend under one pushback |
| Your read on their mission and safety posture | 30s | Honest. Neither performed enthusiasm nor performed skepticism |
| 90-second career narrative | 90s | The through-line, not the résumé |
The falsifier is the move almost nobody makes. Ending "where is AI headed" with "here is what would change my mind" converts an opinion into a position, and it is the clearest available signal that you actually think about this rather than reciting.
Drill Set
| Drill | Cadence | Trains |
|---|---|---|
| Story extraction | Weeks 1–3 | Get the raw material down. Bullets, not prose |
| DTAO rewrite | Weekly | Convert one story to the structure. Decision in sentence one |
| Cold telling, recorded | 2×/week | Random story, 2 min, no notes. Listen back |
| Probe defence | Weekly | I ask the three probes cold. Score the answers |
| The lost-disagreement drill | Biweekly | The hardest story, re-told. It gets better every time |
| Forward-looking rehearsal | Weekly, 15 min | All six questions to a timer |
| Numbers audit | Monthly | Every story must have a measured outcome. Find the ones that do not |
| Anti-over-rehearsal | Monthly | If a recording sounds recited, cut it to bullets and re-derive it live |
Failure Modes
| Failure | Symptom | Fix |
|---|---|---|
| Tour, not decision | Two minutes of context before anything is decided | DTAO: decision in sentence one |
| No disagreement | Every story is frictionless | Pick harder stories. If none have friction, that is a finding |
| Strawmanned opposition | "They just wanted the easy option" | Write their case as they would write it |
| No numbers | "It improved things a lot" | Numbers audit |
| Senior scope | Every story is inside one team | Category 1 is mandatory |
| Feelings-first | Leads with how it felt, or with process | Row 36: the rubric penalizes this explicitly |
| Over-rehearsal | Sounds recited | Cut to bullets and re-derive |
| Under-prepared values round | Improvising on mission and safety | It is reportedly the top failure mode. Written answers, weekly rehearsal |
| Inflated story | Presenting a Senior story as Staff | Do not. Interviewers probe scope and it collapses |
Self-Assessment Rubric
| Level | Standard |
|---|---|
| L0 | Project tours; no decisions; no numbers |
| L1 | Real decisions with tradeoffs; single-team scope; no disagreement |
| L2 | Cross-team decision, named opponent stated fairly, measured outcome |
| L3 | Above, plus alignment achieved through evidence rather than authority; a specific self-critique with a generalizable lesson; and a forward-looking answer that ends on a falsifier |
Hire-bar translation
| Verdict | What it looks like |
|---|---|
| No hire | A tour, no decisions, no disagreement |
| Hire (senior) | Real decisions with tradeoffs; single-team scope |
| Strong hire (senior) | Cross-team decision, named opponent, measured outcome |
| Hire (staff) | Above, plus changed an organization's mind with evidence |
| Strong hire (staff) | Above, plus one decision that was expensive and right, and one that was expensive and wrong |
References
../../research/company-brief.md— mission digest, talking points, the three questions you ask them../../diagnostics/d4-behavioral.md— the diagnostic prompts../../diagnostics/ANSWER-KEY.md— the worked weak/strong contrast../../mocks/README.md— the weekly scored mock- Larson, W. Staff Engineer: Leadership Beyond the Management Track. — the archetypes and what scope means
- Reilly, T. The Staff Engineer's Path. O'Reilly, 2022.
- Fournier, C. The Manager's Path. — useful for the influence-without-authority chapters
- OpenAI. Charter. https://openai.com/charter/
- Anthropic. Core Views on AI Safety. https://www.anthropic.com/news/core-views-on-ai-safety