The report-back contract
A trustworthy agent report-back includes all of the following:- A verdict expressed as a routing action, not a raw ranking.
- A calibration tier for each metric: decision-grade, directional, or noise-band.
- Explicit denominators and coverage, including inconclusive cells.
- Confounds named up front; structural confounds cap claim strength.
- Plainly stated non-results, including when no quality metric is decision-grade.
- Cost with robust centers, such as a median or geomean alongside a mean that may be driven by outliers.
- Native provenance from the freshest Trial Result, not a stale field or a hand-written summary.
- No superiority or rollout claim from tests alone, stale evidence, missing graders, or inspect-only evidence.
Calibration is per metric
Decision-grade means the specific metric has enough valid, complete, calibrated evidence for its declared decision. Directional means the metric is useful for iteration, not promotion. Noise-band means the observed difference cannot support even a directional conclusion. Cross-axis sign consistency can strengthen a noise-band read into a directional hypothesis, but structural confounds permanently cap the language your agent may use. An evidence posture ofdirectional means “iterate, don’t promote.” An
inspect posture blocks rollout and superiority claims until the evidence is
repaired or the claim is narrowed.
Ask these questions next
- Is this decision-grade, directional, or noise-band—and for which metric?
- What is the denominator, and were no-patch, unsure, or unavailable cells kept?
- Were contamination flags verified as true positives, including for the winning arm?
- Is a low or failed rate model signal, or an infrastructure problem?
- Did another rerun supersede this number or caveat?
- Is candidate versus baseline oriented by lineage rather than a directory name?
- Did the summary use the freshest native result rather than a stale field?
- Is the verdict an action I can take now?
A real transparency excerpt
The excerpt below grounds the kind of report-back an agent should produce — 16 matched tasks, 32 arm-task cells. Its verdict is an honest non-answer:inspect here is the fail-closed design working as intended, refusing a
superiority claim the evidence cannot carry rather than manufacturing one.
Its own caveats apply:
- Sanitized excerpt from a real July 2026 comparison.
- The excerpt omits repository paths and arm identity; it is transparency evidence, not a complete Trial Result.
- The inspect verdict blocks superiority and rollout claims.
failed factors say the test cells cannot tell the
two arms apart on this corpus, and the grader evidence lacks the certification
it would need to carry the decision on its own. Grading
explains that certification and the rest of the grader evidence rules.
Example agent report-back
This is a human projection from the excerpt above, not fabricated machine data: “Inspect, low confidence. All 16 matched tasks contributed 32 arm-task cells and all 11 declared graders covered the sample, but the corpus contains no genuine discriminating test-verdict cells. Grader-discrimination certification is absent and incompatible with the protected evidence. Do not make a superiority or rollout claim. Generate the certification before using those grader results.” For a finished evaluation, ask your agent to name the routing action, each metric’s calibration tier, the full denominator and coverage, validity, confounds, observed cost and duration, evidence posture, andnext_action.
The release lifecycle explains how that
report-back can authorize gate, promotion, monitoring, or rollback.