Billing, limits and errors
What evals cost in Executions and in your own provider spend, what is never billed, what happens at the plan limit, the limits of what an eval reads, and every state a result can be in.
Two separate bills
| What it is | Who bills it | |
|---|---|---|
| Units | One per step an eval runs on a trace. A walk that ends at a cheap screen after one step is one unit; one that goes through four steps is four. | Trodo, as Executions on your plan (Settings → Usage) |
| Provider spend | What your model provider charges for judge calls, on your key. | Your provider |
Python, Semantic checks, Filter steps and Human review make no model call and have no provider spend. Analysis shows the two side by side (Provider spend and Units) and never adds them.
This is why triage pays twice: a trace settled by a first-step screen costs one unit and no provider spend; the same trace sent through a judge costs more of both.
Executions are one pool shared with agent steps and alert firings. Settings → Usage splits it into Evals, Agents and Alerts. See Plans and billing for each plan's allowance.
Never billed
- Traces the eval's filters don't match, and traces outside its sample.
- Tests in the editor (a judge in a test still calls your provider).
- A step that couldn't run because of Trodo (the sandbox, or a model it runs, being busy). Those are retried.
When you hit the plan limit
Once the organization has used its Executions for the period, traces aren't scored at all, and are counted under Not scored · over the plan limit until the next period. A walk is never cut off halfway: the decision is made before the first step.
On Pro with Usage beyond your plan turned on (Settings → Billing), and always on Enterprise, evals keep scoring and the extra Executions are billed. See Usage beyond your plan.
Re-scoring history is billed like live traffic, and its estimate is shown before it starts.
Limits
Trace input / output read | 256 KB each (longer arrives truncated, marked) |
| Spans loaded per run | the newest 5,000 |
| Runs loaded per conversation | up to 200: the first and the newest |
| Python run time | 10 seconds per call |
| What a Python step returns / passes on | 256 KB each |
| Judge reasoning kept | 2,000 characters |
| Grounded context | the newest 20 matching spans by default, up to 200 |
| An Intent | 500 characters |
| Re-scoring history | up to 50,000 traces per re-score |
| Audit slice | 0.5% of traffic by default |
States and reasons
Every result says what happened. Passed and failed results reached an ending. The rest:
| State | Reason | What to do |
|---|---|---|
| Waiting | A Human review step | Grade it in the queue |
| Errored | python_exception, python_bad_return, timeout, … | See Python steps → When it errors |
judge_provider_error, judge_no_model, … | See LLM judge → When it errors | |
cross_encoder_missing_input, cross_encoder_config, … | See Semantic checks → When it errors | |
output_type_mismatch, context_mismatch | A step returned something it doesn't declare | |
unit_missing | The trace was deleted before it was scored |
An errored result is never a pass or a fail, and never counts toward a pass rate: evals don't guess.
Traffic that is not scored at all is counted on Analysis under Not scored:
| Reason | Shown as |
|---|---|
filter | did not match the filter |
sampled | outside the sample |
dedup | already scored |
entitlement | over the plan limit |
Who did what
Every eval change, grade, correction and decision is recorded with who made it. Here is where you see it.
Overview
Failures shows everything that broke, grouped by agent, tool and error. Issues are the problems worth fixing: a root cause a coding agent can work from, the traces that prove it, and a history that tells you when a fix did not hold.