Skip to main content
Use this flagship workflow to make a corpus-scoped routing choice, not a claim about models everywhere. Keep the task slice and harness stable enough to attribute a difference.

Say this

What happens

Your agent should smoke the new configuration before investing in a full arm. Then it runs only that new arm on the identical retained slice, freezes the incumbent once as a named baseline, and compares the candidate against that frozen baseline. Have it state expected time and spend before launch. The frozen identity prevents hand-pairing directories or silently changing the reference. If compatible evidence is split across runs, have your agent combine it into one durable result rather than hand-editing anything; combining and reanalyzing existing artifacts spends nothing new. A reasoning-effort comparison is the same shape: hold the model fixed and name the reasoning variant as the treatment. Results stay scoped to the corpus, slice, harness, and time of the evidence. Incomplete, stale, contaminated, or inspect-only evidence cannot support a default or superiority claim.

What a good answer sounds like

The report should name the frozen baseline, the candidate, how many tasks counted, how strong each metric’s evidence is, cost with its outliers, and any confounds — every value from the actual Trial Result. Ask: “Which metric is decision-grade?” “Were all cells retained?” “What does this corpus-scoped result not establish?”