Copy and fill. The artifact that justifies a model choice with evidence, not vibes. Full method: Phase 5.
Author: <name> · Date: <YYYY-MM-DD> · Decision owner: <name> · Re-evaluate by: <date, ~quarterly>
- What the model must do:
<one-sentence task>
- Definition of "good":
<acceptance criteria / target metric value>
- Volume / scale:
<requests/day, tasks/user>
| Constraint | Requirement | Notes |
| Context window | <min tokens> | input + output |
| Max output | <min tokens> | |
| Capabilities | <tool calling? structured output? vision? reasoning?> | |
| Latency ceiling | <p95 target> | under load |
| Data residency / privacy | <region / no-train / self-host?> | may force open-weight |
| Budget | <$ / resolved task> | |
| License (if open) | <commercial OK?> | |
| Model × Provider | Context / max out | Price (in/out $/Mtok) | Capabilities | Notes |
<model@provider> | | | | |
<model@provider> | | | | |
<model@provider> | | | | |
- Golden set:
<size, composition: representative + edge + unanswerable> (Phase 12.01)
- Scoring:
<programmatic / calibrated judge / human>
| Model | Quality (metric) | p95 latency | Cost / resolved task | Safety gate | Notes |
<A> | | | | pass/FAIL | |
<B> | | | | pass/FAIL | |
| Axis | Weight | A | B |
| Quality | <%> | | |
| Cost | <%> | | |
| Latency | <%> | | |
| Reliability | <%> | | |
| Safety | GATE | pass/FAIL | pass/FAIL |
| Weighted score | | | |
- Chosen:
<model@provider> — <why, in 2 sentences>
- Routing:
<easy→cheap model X, hard→premium Y?>
- Fallback:
<provider/model on failure>
- Risks:
<vendor change, price drift, deprecation, quality regression>
- Re-evaluate when:
<quarterly / new model released / cost spike> — re-run this memo on the golden set.
Decision is on the eval, not the catalog. Cards/benchmarks shortlist; your golden set decides.