Say this
What happens
Your agent should turn the onboarding receipt into a narrow evaluation plan: one declared harness change, one task slice, and the existing bounded verifier. Time and cost depend on your slice and provider, so have it state estimates before it starts. Spend draws on the provider account your agent already uses — your existing usage limits, not a separate bill — and a full eval can consume a substantial portion of a $100 or $200 monthly subscription tier’s limits, so treat the pre-launch estimate as a real gate. Docker-backed replay is the normal setup for evidence you intend to act on. The agent should stop for your approval before any model-spending launch and should not quietly broaden the task slice to get more results.While it runs
If you want a check-in, ask for the run’s recorded state rather than a guess:What a good answer sounds like
When the eval finishes, ask for the decision:“Inspect, low confidence. All 16 matched tasks counted and all 11 graders covered them, but the corpus has no genuinely discriminating test cells, so no superiority claim is supported. Next action: generate grader-discrimination certification; no new coding-agent run required.”It names an action, says how many tasks counted, states the caveat plainly, and claims no more than the evidence supports. If a metric is only directional, treat it as a reason to iterate, not to promote. The full report-back contract and the questions to ask next live in Read a Trial Result.