Playbook
Write evals first
The golden set is the product specification. Architecture waits on it.
When: Before choosing LangGraph vs CrewAI, and again every time a prompt changes.
- 01
Collect real tasks
From tickets, traces, and the failures you have already paid for.
Distill Traces into Skills - 02
- 03
Done when
- 20–50 real tasks in CI, with structured asserts, including at least one injection/refusal case.
Refuse
- Eval by demo. 20–50 real tasks before the framework. promptfoo, Inspect, DSPy. CI on prompt change.