Skip to content

Job pack · evalopt

Eval harness

Golden sets, rubrics, CI, red-team. The measurement plane the rest of the field sits on.

When: You need to know whether a prompt, a tool, or a skill got worse.

20–50 real tasks, structured asserts, CI on change, at least one injection/refusal case. Architecture waits on this.

GET /api/canon/jobs/eval-platform?format=md

Matching skill · Eval harness

Playbooks

Doctrine to load

Refuse

  • Eval by demo. Architecture first, golden set never. Success is a recorded GIF.

Checklists

Eval quality

  • 20–50 real tasks, not toys. Include the failures you have already seen.
  • CI on prompt, tool schema, and skill changes.
  • Written rubric. If you cannot score it, you cannot loop it.
  • Generator and evaluator do not share prompt or incentive.

Recipes

Default corpus

promptfooinspect-ailangfusedspy