« Interview Prep · Track Overview

Behavioural & Staff Signal

The half of the loop that is not a whiteboard. Two-in-a-box, disagreement, incident command, regulator conversations, mentorship — and what a senior interviewer is actually listening for.


Table of Contents


1. What they are listening for

At senior/staff/principal level the behavioural round is not a personality check. It is testing four things, and knowing which one a question is aimed at tells you what to include:

SignalThe question behind the questionYou demonstrate it by
ScopeDo you operate at the platform, or at your service?describing consequences for teams you do not own
Judgment under ambiguityWhat do you do when there is no rule?naming the trade-off, deciding, and stating what would change it
Mechanism over heroicsDo you fix the instance or the class?ending stories with what changed in the system, not what you stayed up to do
CalibrationDo you know what you do not know?stating limits before you are pushed on them

The fourth is the one people under-invest in and the one that most distinguishes principal from senior. A candidate who volunteers "here's where this design fails" is read as safe to give authority to; a candidate whose designs have no stated weaknesses is read as not having operated one.

2. The story spine

STAR is fine and insufficient at this level. Use five beats:

  1. Situation — one sentence, with the number that makes it real. "Eleven agents in production, three teams, and no way to tell which agent had done what."
  2. Tension — why it was not obvious. If there is no tension, the story is a task, not a decision.
  3. Decision — what you decided, including the option you rejected and why.
  4. Outcome — with a number, and honestly. Partial successes are more credible than clean ones.
  5. Mechanism — what changed so it cannot recur. This beat is the one that separates levels.

Two minutes. If they want more they will ask, and the follow-up is where the real assessment happens — which means leaving room for it is a tactic, not a compromise.

3. Two-in-a-box

The JD names it explicitly, so expect a direct question. Most candidates treat it as boilerplate; it is a named operating model with specific mechanics.

"What does two-in-a-box mean to you?"

"Undivided accountability, not a partition. The normal EM/PM split says 'you own tech, I own product', and it fails exactly where AI platforms fail — a decision like which autonomy band a new agent gets is both a product decision and a risk decision. Two-in-a-box says both owners are accountable for the same surface: availability, performance, cost, security posture, architectural evolution, the roadmap, and the pager. Including the pager — that one sounds performative and it is the single most effective mechanism in the model, because a product owner who has been woken by a retry storm makes different roadmap decisions without being lobbied."

"How do you keep two people synchronized?"

"Artifacts, and not for process reasons. Either of us can commit the platform in a forum, which requires that we are genuinely synchronized, and two people cannot stay synchronized on shared accountability through conversation alone — one of us will be on holiday when the decision is questioned. So everything material becomes an ADR, an error-budget policy, an ORR record. The test is whether my partner can answer a question about a decision I made without calling me."

"What breaks it?"

"Silence about a disagreement. Two people with goodwill and no protocol resolve disputes by seniority, volume or attrition, and all three are corrosive. So the protocol is agreed in advance, before the first real disagreement, when it is still an abstract conversation."

4. Disagreement

Expect: "Tell me about a time you disagreed with a senior stakeholder."

The move that lands is classifying the disagreement before describing it.

"The first thing I do is ask one question: what evidence would change your mind? If we can both answer it, it is a factual disagreement and we go and measure — that is a good day. If neither of us can answer it, it is a values disagreement and measurement will not resolve it, so the protocol is different: we each write our position, and we escalate both. What we never do is average, because a design at the midpoint of two coherent positions is usually worse than either."

Then a specific story. Structure:

  • what each side wanted, stated fairly — a story where the other person is obviously wrong reads as a story where you were not listening;
  • the falsifier question and what it revealed;
  • what was measured, or what was escalated;
  • whether you were right, and what you did when you were not;
  • disagree-and-commit in writing, so the decision has a record rather than a lingering grudge.

"You lost. Then what?"

"I commit, visibly, and I write down what would make me revisit it. The worst outcome is a half-committed engineer who is quietly right — the platform gets neither the decision that was made nor the one that should have been. And having the revisit condition written down means that if it triggers, reopening is a mechanical thing rather than an I-told-you-so."

5. Incident command

"Walk me through an incident you led."

Structure the answer the way you structured the incident. That is itself the signal:

  1. Detection — how you knew, and how long it took. Time-to-detect is usually the most improvable number and the one people omit.
  2. Roles — incident commander, comms, ops. Named in the first two minutes, before any investigation.
  3. Mitigation — what restored service. Not the fix.
  4. Comms cadence — every 30 minutes, even with nothing new. A commitment, not a courtesy.
  5. Resolution — the actual cause, later, calmly.
  6. Post-mortem — blameless, with named owners and dates.
  7. Mechanism — what changed so the class of incident cannot recur.

Two sentences worth having ready:

"Mitigation is not resolution. Conflating them is how the same incident recurs with a different trigger — you rolled back the bad bundle, and the synchronous fail-closed dependency that turned a bundle bug into a total outage is still there."

"I track post-mortem action-item completion rate as a first-class metric, because everybody writes post-mortems and the completion rate is the only honest measure of whether the culture is a process or a writing exercise."

"What if you don't know what's happening?"

"Mitigate first, understand second. Roll back, shed load, degrade a rung — restore service, then investigate with the pressure off. The instinct to find the cause first is the instinct that turns a twenty-minute incident into a two-hour one."

6. Saying no

A platform role is largely a queue of requests you must decline well. Expect a question about it.

"I try never to say no to a person; I say no to a request, with the reason and the nearest thing I can do. And where I can, I make it a mechanism rather than an opinion — 'the ORR requires a tested rollback and yours is untested' is a conversation about a criterion, and 'I don't think you're ready' is a conversation about me. That is most of what the ORR is for: it turns saying no to a peer from a personality contest into a checklist anybody can apply."

The example to have ready is one where you said no to something urgent and legitimate — not to something obviously bad. Anyone can decline a bad idea. The signal is declining a good idea at the wrong time, and the follow-up you want is "how did that land?"

And the counterpart, which principals volunteer:

"The failure mode of a platform team is becoming the department of no. If the only argument for using the platform is compliance, adoption stalls at exactly the teams with the most leverage. So the standard I hold myself to is that using the platform must be easier than not using it — and where it isn't, that's my bug, not their non-compliance."

7. Regulator and audit conversations

The JD names five forums. Know what each wants — they are different audiences and the same deck fails all five in different ways.

ForumWantsArtifactFails when
Enterprise Architecturefit with the estate, no duplicationreference architecture + integration mapyou present a bespoke stack
Cyberthe threat model and the controlsthreat model, red-team results, OWASP matrixyou say "we have guardrails"
Model Riskvalidation, monitoring, the inventorymodel/agent inventory + validation recordsyou conflate the model with the agent
Internal Auditevidence that controls operatedevidence packs, control-to-evidence mapyour evidence is a screenshot
Group CTTOrisk, cost and roadmap in one pageone pageyou go technical

"How do you talk to a regulator?"

"Plainly, and with an artifact rather than a deck. The framing I use is that the output of an agent run is an evidence pack, not an answer — who authorized it, under which policy version, from what knowledge at what version, which model configuration, what the controls did, who approved, and what happened. If I can hand over that pack for any action they pick, most of the conversation is already answered. And I state the limits myself: the pack proves an artifact is present, not that it is true, which is what independent validation is for."

"What if you don't meet a requirement?"

"Say so, with a plan and a date. Regulators are far more comfortable with a known gap that has an owner than with a surprise found during an examination — and the second one costs you the benefit of the doubt on everything else you said."

8. Mentorship and raising the floor

"How do you raise standards across a team?"

"A standard in code is a control; a standard in a wiki is a suggestion. So my first move is always to find the rule that can move into publish() or into CI. The injection scanner not being on the degradation ladder is a test, not a paragraph in a design doc — the reviewer of that pull request sees a red build with a message instead of having to notice."

"The second is pairing on reviews rather than doing reviews. If I am the only person who catches the missing idempotency key, I am a bottleneck and the floor has not moved. I want the checklist published so a tired reviewer at 5 p.m. still catches it."

"How do you mentor someone more junior than the work requires?"

The three moves from Phase 17, which are concrete and unusual enough to be memorable:

"Give them the demo output before the code, and ask them to find the bug — it teaches reading a system's behaviour rather than its structure. Have them break a seam deliberately and watch exactly one test fail while every component test still passes. And in a chaos exercise, make them predict the degradation before injecting the failure; their prediction is a measurement of their model of the system, and being wrong is the most useful five minutes available."

9. Failure and being wrong

"Tell me about a time you were wrong."

The trap is choosing a failure that is secretly a success. Choose a real one, and make the last beat the mechanism.

What a good answer contains:

  • a decision that was yours, not the team's;
  • the reasoning at the time, stated so it sounds reasonable — because it was, or you would not have made it;
  • what you missed, specifically;
  • the cost, honestly;
  • what changed so that class of mistake is caught by a system rather than by you being smarter next time.

"The pattern I look for in my own mistakes is whether the error was available at the time. If the information was there and I did not look, that is a discipline fix. If it was not there, that is an instrumentation fix — and instrumentation fixes are the ones worth talking about, because they generalize."

10. Influence without authority

For a platform role this is most of the job: you own a system that other teams must adopt and you cannot compel them.

"Three things work, in this order. Make it easier than the alternative — the platform wins on ergonomics or it does not win. Publish the number — 'sixty-three percent of agent actions flow through the platform' moves people and stops 'we're migrating' being a permanent state. And give the stragglers a date and a capability, not a policy — if the only reason to adopt is compliance, the teams with the most leverage will be the last to move, which is exactly backwards."

"A team is bypassing your platform. What do you do?"

"First, find out why — usually it is a capability gap or a latency they cannot afford, and both are my problem. Second, enforce at the action boundary rather than the model boundary: I can live with a team calling a model directly, and I cannot live with a direct path to payments.release. Enforcing where enforcement is cheap and evasion is visible is worth more than a policy everyone agrees to and nobody follows."

11. The five stories to prepare

Write these out. Two minutes each, five beats, ending in a mechanism.

#StoryWhat it demonstrates
1A platform decision that traded something real. Availability for cost, or speed for safety.judgment under ambiguity; you know what you gave up
2An incident you led, with detection time, mitigation, and the mechanism afterwards.scope, calm, mechanism over heroics
3A disagreement with a peer or a senior stakeholder, classified, and how it resolved.the two-in-a-box mechanic, and fairness
4A time you were wrong, with the instrumentation that now catches it.calibration
5Something you made easier for other teams — an adoption story, with the number.influence without authority

One caution: at least one story should be a partial success. A candidate whose every story ends cleanly is a candidate who is telling you about the ones that ended cleanly.

12. Anti-signals

Collected from every phase's staff notes, in rough order of how badly they land:

  • "We" for everything. At this level they need to know what you decided.
  • A story with no tension. That is a task, not a decision.
  • A story with no mechanism. Heroics do not scale and interviewers know it.
  • Describing two-in-a-box as a reporting line. It is an accountability model.
  • Blame in a post-mortem story. Instant, and it is not recoverable.
  • No stated limitation anywhere in the whole conversation. Reads as never having operated one.
  • "I'd escalate" as the answer to every hard question. Escalation is a step, not a decision.
  • A failure story that is secretly a success story. Everyone notices.
  • Answering a behavioural question with architecture. They asked about people; answer about people, then connect it to the mechanism.
  • Not asking anything at the end. For a role defined by shared ownership, having no questions about how that ownership actually works reads as not having thought about the job.