Track F — Behavioral at Staff Altitude

Reported: technical leadership, architecture decisions spanning teams, driving consensus under pressure, concrete tradeoffs rather than soft-skills answers (../../research/source-report.md rows 33–36), plus a recruiter-screen question about where AI is headed (row 4).

And the finding that should change how much time you give this: reported sources name the values/culture round as the leading failure point at a peer lab. For companies with technical bars this high, that is a remarkable claim — and senior engineers systematically under-prepare here because it does not feel like real work.

→ Study guide: WARMUP.md — all twelve story categories with worked model answers at staff density, the probe playbook, and the forward-looking answers written out in full.


Table of Contents


Why This Is Graded at Staff, Not Senior

AI-lab levelling is compressed. Reported consensus: an OpenAI level maps roughly one level higher than the same number at Google or Meta, and L5 — titled "Senior" — carries Staff-equivalent scope (../../research/findings.md).

The practical consequence for every story you tell:

Senior storyStaff story
"I designed and built X""I decided X over Y, and got three teams to agree"
"It improved latency 30%""It improved latency 30% and cost the storage team 60% more memory, which is why they pushed back"
"I convinced them""I built a prototype that made the argument for me"
"It went well""It went well; here is the part I got wrong and what it taught me"
One teamCross-team, with a named opponent

The reliable test: if nobody in the story disagreed with you, it is not a Staff story. Consequential decisions attract disagreement. A story with no opposition is either not consequential or you have edited the opposition out — and interviewers probe for exactly this.


DTAO, Not STAR

STAR front-loads Situation — the least interesting part — and buries the decision. It was designed for interviews that wanted to know whether you were a good teammate. This round wants to know how you think.

LetterSectionShare
D — DecisionWhat you decided, in one sentence, first1 sentence
T — TradeoffThe alternatives and why each lost. Numbers live here~40%
A — AlignmentWho disagreed, and what you actually did about it~30%
O — OutcomeMeasured, including what you got wrong~25%

Context goes in a clause, not a paragraph: "On the multilingual ranking pipeline, we decided X." That is enough. If they need more they will ask — and them asking is a good sign, because it means they are engaged rather than waiting for you to finish.


The Story Bank

Twelve to fifteen stories from your actual history. I will not invent them, embellish them, or let a Senior-scope story be presented as Staff-scope.

The raw material I need from you, per story — bullets are fine, prose is not required:

  1. What was the decision? (one sentence)
  2. What constraint made it hard?
  3. What alternatives did you seriously consider?
  4. Who disagreed, and what did they want instead?
  5. What did you do to get alignment?
  6. What was the measured outcome?
  7. What did you get wrong?

Your history — multilingual search and recommendation, media streaming, networking, enterprise infra, cloud — is unusually rich for this. Ranking pipeline redesigns, index-serving migrations, embedding infrastructure decisions, and cross-org platform migrations are all naturally Staff-altitude if you write the tradeoff rather than the tour.

Each story lands in stories/NN-slug.md with: the DTAO write-up, a 90-second spoken version, a 2-minute version, the tags it covers, and its probe list.


Required Story Categories

Every one must be filled. A gap here is a gap the interviewer will find.

#CategoryWhy it is askedStatus
1Architecture decision affecting multiple teamsRow 34 — the core Staff signal
2A disagreement you lostThe highest-signal prompt that exists. See below
3A disagreement you won, and why they concededTests whether you persuade or just outlast
4An outage you ownedOwnership under pressure; blameless analysis
5A project you killed or descopedSunk-cost resistance. Rare and valuable
6Mentoring / raising a team's barScope beyond your own output
7A bet that failedCalibrated risk-taking, honestly reported
8Driving consensus without authorityRow 35
9Shipping under a hard deadline with quality tensionThe tradeoff nobody escapes
10A time you changed your mind from dataUpdates on evidence
11Working with non-engineering partnersReported: collaboration with researchers, PMs, safety
12Something you built that you would now build differentlyTechnical judgement over time

Category 2 deserves its own note

Four ways it fails, all of them visible from across the room:

FailureWhat it sounds likeWhat it signals
Humble-brag"I lost, but six months later they did it my way"You cannot actually update
Victim"Management overruled me for political reasons"You do not distinguish wrong from outvoted
TrivialA disagreement about namingYou avoid consequential conflict
Revisionist"In hindsight they were right about everything"Performed humility; no real position

What works: state your position as strongly as you actually held it, state theirs fairly enough that they would recognize it, say what decided it and whether the process was sound even if the outcome was not, and say how you behaved after losing — committed or sandbagged.

Then give your honest current read. All three of these are strong:

  • "They were right, and here is what I had not weighted properly."
  • "I still think I was right, and here is the evidence that has accumulated since."
  • "We were both solving the wrong problem."

Only performed humility is weak.


Probe Lists

Every story needs three follow-ups written out — the questions an interviewer asks to test whether the story is real. Generic examples; each story gets its own specific set:

  1. "What was their strongest argument?" — the single most discriminating probe in the round. If you cannot produce a strong version of the opposing case, you never engaged with it, and the whole story becomes suspect.
  2. "What would have had to be true for the other option to win?" — tests whether you modelled the decision or pattern-matched it.
  3. "Who else was affected that you did not mention?" — tests scope honesty.
  4. "What did that cost the other team?" — cross-team decisions always cost someone.
  5. "How long did it take, and how much of that was the disagreement?" — tests whether the consensus story is real.
  6. "What did you measure, and how did you know it was not a coincidence?" — tests rigour.
  7. "What would you do differently?" — the answer must be specific, not "communicate more."

The Forward-Looking Questions

Written answers, rehearsed weekly, kept in forward/. Company-specific material lives in ../../research/company-brief.md.

QuestionLengthThe bar
Where is AI headed?90sA specific falsifiable claim + evidence + a falsifier + what you would build
Why this company?30sGrounded in the work, not the brand
What would you work on?30sConcrete, and connected to what you have shipped
What is their hardest unsolved engineering problem?60sA real technical position you can defend under one pushback
Your read on their mission and safety posture30sHonest. Neither performed enthusiasm nor performed skepticism
90-second career narrative90sThe through-line, not the résumé

The falsifier is the move almost nobody makes. Ending "where is AI headed" with "here is what would change my mind" converts an opinion into a position, and it is the clearest available signal that you actually think about this rather than reciting.


Drill Set

DrillCadenceTrains
Story extractionWeeks 1–3Get the raw material down. Bullets, not prose
DTAO rewriteWeeklyConvert one story to the structure. Decision in sentence one
Cold telling, recorded2×/weekRandom story, 2 min, no notes. Listen back
Probe defenceWeeklyI ask the three probes cold. Score the answers
The lost-disagreement drillBiweeklyThe hardest story, re-told. It gets better every time
Forward-looking rehearsalWeekly, 15 minAll six questions to a timer
Numbers auditMonthlyEvery story must have a measured outcome. Find the ones that do not
Anti-over-rehearsalMonthlyIf a recording sounds recited, cut it to bullets and re-derive it live

Failure Modes

FailureSymptomFix
Tour, not decisionTwo minutes of context before anything is decidedDTAO: decision in sentence one
No disagreementEvery story is frictionlessPick harder stories. If none have friction, that is a finding
Strawmanned opposition"They just wanted the easy option"Write their case as they would write it
No numbers"It improved things a lot"Numbers audit
Senior scopeEvery story is inside one teamCategory 1 is mandatory
Feelings-firstLeads with how it felt, or with processRow 36: the rubric penalizes this explicitly
Over-rehearsalSounds recitedCut to bullets and re-derive
Under-prepared values roundImprovising on mission and safetyIt is reportedly the top failure mode. Written answers, weekly rehearsal
Inflated storyPresenting a Senior story as StaffDo not. Interviewers probe scope and it collapses

Self-Assessment Rubric

LevelStandard
L0Project tours; no decisions; no numbers
L1Real decisions with tradeoffs; single-team scope; no disagreement
L2Cross-team decision, named opponent stated fairly, measured outcome
L3Above, plus alignment achieved through evidence rather than authority; a specific self-critique with a generalizable lesson; and a forward-looking answer that ends on a falsifier

Hire-bar translation

VerdictWhat it looks like
No hireA tour, no decisions, no disagreement
Hire (senior)Real decisions with tradeoffs; single-team scope
Strong hire (senior)Cross-team decision, named opponent, measured outcome
Hire (staff)Above, plus changed an organization's mind with evidence
Strong hire (staff)Above, plus one decision that was expensive and right, and one that was expensive and wrong

References