Product Discovery
Phase 15 · Document 01 · Startup Playbook Prev: 00 — Startup Opportunity Map · Up: Phase 15 Index
Table of Contents
- Why This Matters
- Core Concept
- Mental Model
- Hitchhiker's Guide
- Warmup Readings
- Deep Readings and External References
- Key Terms
- Important Facts
- Observations from Real Systems
- Common Misconceptions
- Engineering Decision Framework
- Hands-On Lab
- Verification Questions
- Takeaways
- Artifact Checklist
1. Why This Matters
The most common way an AI startup dies is building something nobody needs — beautifully engineered, technically impressive, and unwanted. Engineers are especially prone to this: the LLM makes it so easy to build an impressive demo that you skip the only thing that matters — confirming a real, acute, expensive problem that someone will pay to solve. Product discovery is the disciplined practice of finding and validating that problem before you build. It's the cheapest, highest-leverage work in the entire playbook: a week of talking to users can save a year of building the wrong thing. For LLM products specifically, discovery has a twist — the demo magic ("wow, it summarized my contract!") fools both you and early users into thinking you have a product when you have a party trick. This doc is how to cut through that and find a problem worth a company.
2. Core Concept
Plain-English primer: fall in love with the problem, not your solution
Product discovery means deeply understanding a specific customer's specific problem — before committing to a solution. The failure mode is solution-first ("I'll build an AI agent for X") instead of problem-first ("what is the most painful, expensive, frequent thing this person does, and would they pay to make it go away?"). You validate by talking to real potential customers and watching what they actually do and pay for — not what they say they like.
SOLUTION-FIRST (the trap): "I built an AI tool for lawyers" → search for someone who wants it → usually nobody pays
PROBLEM-FIRST (the way): talk to lawyers → find the acute, expensive, frequent pain → build the smallest thing that kills it → they pay
Painkiller vs vitamin (the single most important filter)
- A vitamin is nice-to-have — "this is cool, I might use it." People rarely pay for vitamins, and they churn.
- A painkiller is must-have — it removes acute, expensive, recurring pain. People pay for painkillers and keep paying.
The best LLM opportunities are painkillers for expensive knowledge work: tasks that today cost a lot of skilled human time (legal review, clinical documentation, support, code, compliance, 00). The discovery question isn't "would this be useful?" (everything is vaguely useful) — it's "is this pain acute and expensive enough that they'll pay and switch?"
Jobs To Be Done (JTBD): what are they "hiring" the product to do?
The JTBD lens reframes discovery: customers don't want your product, they want a job done. "I'm not hiring a drill, I'm hiring a hole." For an LLM product: a lawyer isn't hiring "an AI" — they're hiring "get this 80-page contract reviewed for risky clauses in 10 minutes instead of 3 hours." Discover the job (the outcome, the context, the current way they get it done, the frustrations) and you'll know what to build and how to pitch it.
The validation interview (and its traps)
The core tool is the customer interview, done well (per The Mom Test): ask about their life and past behavior, not your idea.
- Good: "Walk me through the last time you did X. How long did it take? What was annoying? What did you do about it? What did it cost?"
- Bad: "Would you use an AI tool that does X?" (everyone says yes to be nice — worthless data).
The rule: talk about their problem, not your solution; ask about the past, not the hypothetical future; let them talk. Compliments are not validation; commitments are (time, money, a real intro, a signed pilot).
Signals of real demand (vs polite interest)
WEAK (polite): "cool!", "I'd totally use that", "let me know when it launches", lots of feature requests
STRONG (real): "how much is it?", "can I use it now?", they already hacked a workaround, they introduce you to their boss,
they give you their data, they pre-pay / sign a pilot, they're annoyed it doesn't exist yet
The crispest test: "Would you pay $X/month for this?" beats "Would you use this for free?" — willingness to pay is the truth serum. Even better is an actual commitment (a pilot, a deposit, a contract).
The discovery loop (talk → demo → measure → decide)
Discovery is iterative and fast:
- Pick a narrow segment and a hypothesized pain.
- Interview 10+ people about how they do that job today (problem, not solution).
- Build the smallest possible demo (hardcoded prompts, a CLI/Streamlit — see 04) to make the pain tangible.
- Show it and observe reactions and commitments (not compliments).
- Decide: strong demand → keep going; weak → pivot (new pain/segment) or kill. Discovery's job is to kill bad ideas cheaply — a fast "no" is a win.
The AI-specific discovery trap: demo magic
LLMs make impressive demos trivial, which corrupts discovery in two ways: (1) you fall in love with the demo and skip validation; (2) users are wowed by the novelty and give false positive signals that don't convert to payment. Counter it by pushing past the wow to "would you pay, switch your workflow, and trust it on your real data daily?" A demo that impresses but doesn't get a commitment is a vitamin in disguise. Also probe trust/accuracy requirements early — for many valuable jobs, "usually right" isn't good enough, and that shapes the whole product (Phase 12/Phase 14).
3. Mental Model
FALL IN LOVE WITH THE PROBLEM, NOT YOUR SOLUTION. #1 startup killer = building what nobody needs.
SOLUTION-FIRST (trap) vs PROBLEM-FIRST (way): talk to users → find acute/expensive/frequent pain → smallest thing that kills it → they pay
★ PAINKILLER vs VITAMIN: people pay for painkillers (acute/expensive/recurring — esp. expensive KNOWLEDGE WORK [00]), not vitamins.
the question is NOT "useful?" but "acute+expensive enough to PAY and SWITCH?"
JTBD: customers hire the product to get a JOB done (the outcome) — discover the job, its context, current way, frustrations.
INTERVIEW (Mom Test): ask about their LIFE + PAST behavior, NOT your idea.
good: "walk me through the last time you did X — how long, what was annoying, what did it cost?"
bad: "would you use an AI that does X?" (everyone says yes — worthless)
SIGNALS: weak=compliments/feature-requests/"let me know when it launches" · STRONG=commitments (money, "can I use it now?", intro to boss, their data, signed pilot, pre-pay)
truth serum: "would you pay $X/mo?" > "would you use it free?"
LOOP: segment+pain → interview 10+ → tiny demo [04] → observe reactions/COMMITMENTS → DECIDE (strong→go, weak→pivot/KILL). killing bad ideas cheap = a WIN.
★ AI TRAP — DEMO MAGIC: LLM demos wow easily → false positives. push past wow to "pay + switch + TRUST on real data daily?" (accuracy/trust reqs shape the product [12/14])
Mnemonic: find an acute, expensive pain (a painkiller, not a vitamin) by talking to real users about their past behavior — not pitching your idea — and trust commitments (money, pilots, their data), not compliments. LLM demo magic creates false positives, so push past the wow to "would you pay, switch, and trust it daily?"
4. Hitchhiker's Guide
What to look for first: is the pain acute, expensive, and frequent — a painkiller? And will they commit (pay, pilot, give data), not just compliment? Those two signals are the whole game.
What to ignore at first: building anything beyond a throwaway demo, feature requests, and praise. Discovery is about learning, not building or being liked.
What misleads beginners:
- Solution-first thinking. "I'll build an AI for X" before knowing if X hurts — start with the pain.
- Pitching in interviews. Asking "would you use my idea?" gets polite yeses — ask about their past behavior (Mom Test).
- Mistaking compliments for validation. Only commitments count (money/pilot/data/intro).
- Demo magic. LLM demos wow easily — push to "would you pay and switch?" (vitamin-in-disguise).
- Building a vitamin. Cool but non-acute pain won't sustain payment or beat the status quo.
- Skipping the trust question. For valuable jobs, "usually right" may be unacceptable — discover the accuracy bar (Phase 12).
How experts reason: they fall in love with a specific painful job (JTBD), interview 10+ real users about past behavior (not the idea), look for painkiller-grade pain in expensive knowledge work, treat commitments as the only real signal, use a tiny demo to provoke reactions (not to ship), push past LLM demo magic to willingness-to-pay/switch/trust, and kill or pivot fast when demand is weak. They know a cheap "no" is a successful discovery.
What matters in production (of discovery): evidence of acute, expensive pain; real commitments from the target segment; a clear JTBD; and a validated (or killed) hypothesis — before meaningful engineering.
How to debug/verify: count commitments vs compliments across interviews; if all you have is praise and feature ideas, you haven't validated. Test "would you pay $X?" explicitly; if the demo wows but no one commits, you have a vitamin.
Questions to ask: is this a painkiller or vitamin? how expensive/frequent is the pain? what's the JTBD? did I ask about past behavior or pitch my idea? did anyone commit (pay/pilot/data)? is the demo-magic fooling me? what's the trust/accuracy bar?
What silently wastes a year: solution-first building, pitching instead of listening, counting compliments as validation, demo magic, and refusing to kill a vitamin.
5. Warmup Readings
| Title | Why to read it | What to extract | Difficulty | Time |
|---|---|---|---|---|
| 00 — Startup Opportunity Map | Where pain is worth most | painkiller categories | Beginner | 25 min |
| 02 — Market Selection | Who exactly to talk to | ICP, beachhead | Beginner | 25 min |
| 04 — MVP Design | The tiny validation demo | smallest valuable thing | Beginner | 25 min |
| Phase 12 — Evaluation | The trust/accuracy bar | "usually right" isn't enough | Beginner | 20 min |
6. Deep Readings and External References
| Title | URL | Why it matters | Read first | Lab connection |
|---|---|---|---|---|
| The Mom Test (Rob Fitzpatrick) | https://www.momtestbook.com/ | How to interview without lying to yourself | past behavior, not ideas | This lab |
| Clayton Christensen — Jobs To Be Done | https://hbr.org/2016/09/know-your-customers-jobs-to-be-done | The JTBD lens | hire-a-product framing | This lab |
| YC — How to Talk to Users (Eric Migicovsky) | https://www.ycombinator.com/library/6g-how-to-talk-to-users | Practical interview tactics | good vs bad questions | This lab |
| YC — How to Get Startup Ideas (Paul Graham) | https://www.paulgraham.com/startupideas.html | Finding real problems | live in the future / notice pain | Concept |
| Superhuman PMF engine (Rahul Vohra) | https://review.firstround.com/how-superhuman-built-an-engine-to-find-product-market-fit/ | Measuring demand | "very disappointed" test | Validation |
7. Key Terms
| Term | Simple meaning | Technical meaning | Why it matters | Where it appears | How to use it |
|---|---|---|---|---|---|
| Product discovery | Find the real problem | Validate pain before building | Avoids building unwanted | this doc | Do it first |
| Problem-first | Start from pain | Pain → solution, not reverse | Prevents solution bias | this doc | Reframe |
| Painkiller | Must-have | Acute/expensive/recurring need | People pay | filter | Target it |
| Vitamin | Nice-to-have | Non-acute, optional | People don't pay | filter | Avoid |
| JTBD | The job hired | Outcome + context customer wants | Reveals real need | discovery | Discover the job |
| Mom Test | Honest interviewing | Ask past behavior, not idea | Avoids false positives | interviews | Apply it |
| Commitment | Real signal | Money/pilot/data/intro | True validation | signals | Seek it |
| Demo magic | LLM wow effect | Impressive but non-converting | False positives | AI trap | Push past it |
8. Important Facts
- The #1 startup killer is building something nobody needs — discovery (validating the pain first) is the cheapest, highest-leverage work.
- Be problem-first, not solution-first — start from an acute pain, then build the smallest thing that kills it.
- Painkiller vs vitamin is the key filter — people pay for painkillers (acute/expensive/recurring), especially in expensive knowledge work (00); they don't pay for vitamins.
- JTBD: customers hire the product to get a job done — discover the outcome, context, and frustrations, not feature wishes.
- Interview about past behavior, not your idea (the Mom Test) — "would you use my AI?" gets worthless polite yeses.
- Commitments (money, pilots, data, intros) are the only real validation — compliments and feature requests are not.
- "Would you pay $X/month?" beats "would you use it for free?" — willingness to pay is the truth serum.
- LLM demo magic creates false positives — push past the wow to "would you pay, switch, and trust it on real data daily?"; discover the accuracy/trust bar early (Phase 12/Phase 14).
9. Observations from Real Systems
- "ChatGPT for X" graveyards are full of impressive demos that never found a paying, switching customer — demo magic without painkiller validation.
- The winners did deep discovery in a narrow segment — Cursor (developers' real editing pain), Harvey (lawyers' document workload), Abridge (clinicians' note-taking burden) — they validated acute, expensive pain before scaling (00).
- The Mom Test / "talk to users" is YC's most-repeated advice because founders so reliably pitch instead of listen and hear false positives.
- Willingness-to-pay flips many "validated" ideas — users who loved the free demo vanish at "$50/month," exposing vitamins.
- Accuracy/trust requirements reshape products — discovery often reveals the job needs near-perfect reliability or human-in-the-loop, changing scope and pricing (Phase 10.05).
10. Common Misconceptions
| Misconception | Reality |
|---|---|
| "If I build it impressive, they'll come" | Most impressive demos serve no acute need |
| "Users said they'd use it — validated!" | Compliments aren't commitments; ask for payment |
| "Asking about my idea is good research" | Pitching biases answers — ask about past behavior |
| "Any useful tool will sell" | Vitamins don't sell; only painkillers do |
| "The wow factor means we have a product" | Demo magic creates false positives — push to pay/switch |
| "More features will create demand" | Demand comes from acute pain, not feature count |
11. Engineering Decision Framework
PRODUCT DISCOVERY (validate the pain before building):
1. PICK a narrow segment + a hypothesized PAINFUL JOB (problem-first, from [00]/[02]).
2. INTERVIEW 10+ real users about PAST BEHAVIOR (Mom Test): how do they do this job today, time, cost, frustrations? (don't pitch)
3. CLASSIFY the pain: PAINKILLER (acute/expensive/recurring → continue) or VITAMIN (kill/pivot)?
4. TINY DEMO [04]: smallest possible thing (hardcoded prompts/CLI) to make the pain tangible.
5. MEASURE COMMITMENTS not compliments: "would you pay $X?", "can I use it now?", pilot, data, intro. Push PAST demo magic.
6. PROBE trust/accuracy bar [12/14] — does "usually right" work for this job?
7. DECIDE: strong demand → carry to market selection [02] + MVP [04]; weak → PIVOT (new pain/segment) or KILL (cheap no = win).
| Signal | Action |
|---|---|
| Acute, expensive, frequent pain + commitments | Proceed (painkiller) |
| Wow but no willingness to pay | Vitamin — pivot/kill |
| Compliments + feature requests only | Keep interviewing; not validated |
| "Usually right" is unacceptable for the job | Redesign for trust/HITL [10.05] |
| Nobody describes a current workaround | Pain likely not acute |
12. Hands-On Lab
Goal
Run a one-week discovery sprint on the category chosen in 00: interview users about a painful job, build a tiny demo, and reach a go/pivot/kill decision based on commitments.
Prerequisites
- The chosen category/segment from 00; a list of 10 reachable potential users; a way to build a hardcoded-prompt demo (CLI/Streamlit, 04).
Steps
- Hypothesize the painful job: write the specific job and segment in one sentence (JTBD framing).
- Interview 10+ (Mom Test): ask how they do this job today — time, cost, frustrations, current workarounds — without pitching. Log quotes.
- Classify the pain: painkiller or vitamin? Acute/expensive/frequent? Note evidence.
- Build a tiny demo: the smallest hardcoded-prompt version that makes the pain tangible (04).
- Provoke commitments: show 5 people; ask "would you pay $X/month?", "can I pilot with your real data?", "intro me to whoever owns this?" — record commitments vs compliments, and push past the demo wow.
- Decide: strong commitments → go (to 02/04); weak → pivot (new pain/segment) or kill.
Expected output
A discovery report: the hypothesized job, 10+ interview notes (past-behavior based), a painkiller/vitamin verdict, a tiny demo, a tally of commitments vs compliments, and a justified go/pivot/kill decision.
Debugging tips
- All praise, no commitments → you pitched instead of probed, or it's a vitamin.
- Demo wowed but "$X?" got nos → demo magic; the pain isn't acute/expensive enough.
Extension task
Run the Superhuman PMF test ("how would you feel if you could no longer use this?") on early users; >40% "very disappointed" is a strong signal.
Production extension
Carry a validated pain into market selection (02) and MVP design (04); fold discovered accuracy/trust requirements into the eval plan (Phase 12).
What to measure
Pain acuteness, commitments vs compliments ratio, willingness-to-pay, demo-to-commitment conversion, go/pivot/kill outcome.
Deliverables
- A JTBD statement + 10+ Mom-Test interview notes.
- A painkiller/vitamin verdict with evidence.
- A tiny demo + a commitments tally + a go/pivot/kill decision.
13. Verification Questions
Basic
- What is product discovery and why is it the highest-leverage early work?
- What's the difference between a painkiller and a vitamin?
- What's the core rule of the Mom Test?
Applied 4. Why are commitments better validation than compliments? 5. What is "demo magic" and how does it corrupt LLM product discovery?
Debugging 6. Every interview loves your idea but no one will pay. What happened? 7. Your demo wows but conversion is zero. What does that tell you?
System design 8. Design a one-week discovery sprint for a vertical AI agent idea.
Startup / product 9. How do you decide go/pivot/kill from discovery evidence, and why is a fast "no" a win?
14. Takeaways
- Fall in love with the problem, not your solution — the #1 startup killer is building what nobody needs.
- Filter for painkillers, not vitamins — acute, expensive, recurring pain (especially expensive knowledge work) is what people pay for (00).
- Interview about past behavior (Mom Test), not your idea — and trust commitments (money/pilots/data), never compliments.
- Push past LLM demo magic to "would you pay, switch, and trust it daily?" — and discover the accuracy/trust bar (Phase 12/Phase 14).
- Discovery's job is to kill bad ideas cheaply — a fast, well-evidenced "no" beats a year building the wrong thing.
15. Artifact Checklist
- A JTBD statement (specific job + segment).
- 10+ Mom-Test interviews (past behavior, logged quotes).
- A painkiller/vitamin verdict with evidence.
- A tiny demo + a commitments-vs-compliments tally + willingness-to-pay.
- A justified go / pivot / kill decision.
Up: Phase 15 Index · Next: 02 — Market Selection