Skip to main content
Your first eval turns a credible onboarding slice into one bounded measurement. You set the question and the spend boundary; your agent proposes a plan, runs it after you confirm, and reads the Trial Result for you.

Say this

What happens

Your agent should turn the onboarding receipt into a narrow evaluation plan: one declared harness change, one task slice, and the existing bounded verifier. Time and cost depend on your slice and provider, so have it state estimates before it starts. Spend draws on the provider account your agent already uses — your existing usage limits, not a separate bill — and a full eval can consume a substantial portion of a $100 or $200 monthly subscription tier’s limits, so treat the pre-launch estimate as a real gate. Docker-backed replay is the normal setup for evidence you intend to act on. The agent should stop for your approval before any model-spending launch and should not quietly broaden the task slice to get more results.

While it runs

If you want a check-in, ask for the run’s recorded state rather than a guess:
If the run looks stuck or broken, use Troubleshooting — it covers how your agent should classify and recover each failure state.

What a good answer sounds like

When the eval finishes, ask for the decision:
A good report-back reads like this, with every value taken from the actual Trial Result:
“Inspect, low confidence. All 16 matched tasks counted and all 11 graders covered them, but the corpus has no genuinely discriminating test cells, so no superiority claim is supported. Next action: generate grader-discrimination certification; no new coding-agent run required.”
It names an action, says how many tasks counted, states the caveat plainly, and claims no more than the evidence supports. If a metric is only directional, treat it as a reason to iterate, not to promote. The full report-back contract and the questions to ask next live in Read a Trial Result.