Say this
What happens
Your agent should smoke the new configuration before investing in a full arm. Then it runs only that new arm on the identical retained slice, freezes the incumbent once as a named baseline, and compares the candidate against that frozen baseline. Have it state expected time and spend before launch. The frozen identity prevents hand-pairing directories or silently changing the reference. If compatible evidence is split across runs, have your agent combine it into one durable result rather than hand-editing anything; combining and reanalyzing existing artifacts spends nothing new. A reasoning-effort comparison is the same shape: hold the model fixed and name the reasoning variant as the treatment.
Results stay scoped to the corpus, slice, harness, and time of the evidence.
Incomplete, stale, contaminated, or inspect-only evidence cannot support a
default or superiority claim.