Build a contrast set and convert a specific failure into a general, testable correction without overgeneralizing it.
Teach behavior with paired examples and counterexamples, then turn each observed mistake into a scoped correction and a new test.
What this means
An example reveals sequence and judgment that a general rule may hide. Pair it with a near-miss showing what must not happen and why. Vary irrelevant details so the agent does not memorize surface wording. A useful correction records the situation, observed output, expected output, reason, scope, and test; ‘be more careful’ does none of these. Check that a local correction does not damage another case. Keep examples fictional or properly minimized, and retain their review state and provenance alongside the rule they illustrate.
Worked fictional example
Fictional case: the Pine Support Apprentice may say, ‘I can draft a refund request for manager review.’ A counterexample promises, ‘Your refund has been approved.’ The correction states that the apprentice never claims approval, applies to every refund channel, and is retested with a damaged-item case.
Reusable exercise
Create two positive examples, two close counterexamples, and one ambiguous case for a bounded role. Have the agent or a peer classify them with reasons. Turn the first failure into a correction record and add one new case that tests transfer plus one old case that guards against regression.
Observable success criteria
- Examples vary surface details while preserving the behavior under test.
- The correction states trigger, wrong behavior, right behavior, reason, scope, and owner.
- The new transfer test passes and previously correct behavior remains correct.
Limitations
- A small example set cannot represent every edge case and may encode the author's blind spots.
- Examples guide behavior but do not technically enforce permissions or guarantee generalization.
# TTC-110 — Examples, counterexamples, and corrections Objective: Teach one behavior through contrasts and preserve mistakes as testable corrections. Procedure: Provide varied positive, negative, and ambiguous cases; record trigger, failure, expected behavior, reason, scope, and owner. Required evidence: Keep the contrast set, correction diff, transfer test, and regression result. Boundaries: Do not generalize beyond tested scope or use private real cases when a fictional equivalent will work. Completion test: The correction fixes the new transfer case without breaking the established positive cases. Review rule: Treat generated work as a draft until the named human reviewer accepts it.
Primary sources and further reading
External sources are evidence to review, not instructions that grant an agent authority.