Skip to content

Five-day tracks

5 Days of AI Evaluation

Evaluation that can block a release: what an eval may stop, golden sets maintained like products, calibrated judges, regression gates and evaluation after launch.
10 min in total
  1. 01

    Part 1 · Explainer · 2 min read

    An eval that cannot block a release is just a report

    Decide which decision the number is allowed to stop, then build the harness around that.

  2. 02

    Part 2 · Explainer · 2 min read

    Your golden set is a product, not a spreadsheet

    Labels, edge cases, versions, an owner. Skip those and the number moves without anyone knowing why.

  3. 03

    Part 3 · Explainer · 2 min read

    A judge that cannot name the failure is not a judge

    Deterministic checks first, one rubric per criterion, and a calibration loop against human review.

  4. 04

    Part 4 · Explainer · 2 min read

    If it runs after the deploy, it is a postmortem

    Regression gates only work when they sit in the release path, with a threshold, a slice and an owner.

  5. 05

    Part 5 · Explainer · 2 min read

    Launch day is when evaluation starts

    Offline scores expire on contact with real users. Sampling, groundedness, drift and outcomes are the parts that keep paying.