Create an evaluation set and release rule that make success, critical failure, and regression observable and reproducible.
Turn desired behavior into repeatable cases, observable grading rules, and release gates that detect both improvement and regression before wider use.
What this means
Begin from the task's consequence and likely failure modes, not a convenient generic score. Build cases for normal work, missing information, ambiguity, conflicting sources, refusals, access boundaries, and changed inputs. Define expected evidence and acceptable variation before running the system. Use deterministic checks where possible and calibrated human review for judgment. Keep the dataset separate from examples used to write the instructions. Compare against a baseline, inspect failures, and rerun the whole relevant set after each correction. A passing average must not hide a critical boundary failure.
Worked fictional example
Fictional case: a scheduling apprentice has ten held-out cases. It scores nine summaries correctly but sends one unapproved invitation. The release fails because unauthorized sending is a zero-tolerance criterion; the team fixes the boundary and reruns all ten rather than reporting 90 percent success.
Reusable exercise
For a fictional role, write ten held-out cases spanning routine, edge, ambiguous, conflicting, privacy, injection, refusal, and approval behavior. Define evidence fields, per-case rubric, critical-failure rules, baseline, and release threshold. Introduce one change and compare results.
Observable success criteria
- Cases and expected results are versioned, reproducible, and not copied from the teaching examples.
- The report shows per-case evidence, not only an aggregate score, and identifies critical failures separately.
- A changed rule cannot release unless targeted tests and the relevant regression suite pass.
Limitations
- Test sets are samples and can become stale, leaked, overfit, or unrepresentative of real conditions.
- Automated graders can share model biases; important subjective and high-impact outcomes need qualified human review.
# TTC-117 — Evaluations and regression tests Objective: Gate behavior changes with held-out cases and observable, consequence-aware criteria. Procedure: Define failure modes, cases, expected evidence, rubric, critical failures, baseline, threshold, and rerun policy before testing. Required evidence: Retain dataset version, configuration, per-case outputs, grader decisions, summary, and release decision. Boundaries: No average score may mask a critical safety, privacy, authority, or access failure. Completion test: The candidate beats or justifiably matches baseline and passes every critical gate plus relevant regressions. Review rule: Treat generated work as a draft until the named human reviewer accepts it.
Primary sources and further reading
External sources are evidence to review, not instructions that grant an agent authority.