Skip to main content
Use this workflow when an instruction-file change has one testable behavior hypothesis. Keep the model, task slice, and other harness settings fixed.

Say this

What happens

Your agent should first propose the baseline, candidate, retained slice, verifier, graders, and expected time and spend. A smoke run can expose setup or obvious behavior problems, but it is not a keep decision. A credible instruction-file signal uses at least 10 retained tasks and isolates one lever at a time. For a quality claim, ask for the recommended dimensions: clarity, simplicity, coherence, intentionality, robustness, instruction_adherence, scope_discipline, and diff_minimality. Missing coverage or invalid replay makes the result diagnostic rather than a rollout decision.

What a good answer sounds like

The report should name the changed lever, how many matched tasks counted, how strong the evidence is for each metric, and any confounds — every value from the actual Trial Result. Ask: “Which metrics are decision-grade?” “Were all matched cells retained?” “What blocks a keep decision?”