Billing, limits and errors

What evals cost in Executions and in your own provider spend, what is never billed, what happens at the plan limit, the limits of what an eval reads, and every state a result can be in.

Two separate bills

What it isWho bills it
UnitsOne per step an eval runs on a trace. A walk that ends at a cheap screen after one step is one unit; one that goes through four steps is four.Trodo, as Executions on your plan (Settings → Usage)
Provider spendWhat your model provider charges for judge calls, on your key.Your provider

Python, Semantic checks, Filter steps and Human review make no model call and have no provider spend. Analysis shows the two side by side (Provider spend and Units) and never adds them.

This is why triage pays twice: a trace settled by a first-step screen costs one unit and no provider spend; the same trace sent through a judge costs more of both.

Executions are one pool shared with agent steps and alert firings. Settings → Usage splits it into Evals, Agents and Alerts. See Plans and billing for each plan's allowance.

Never billed

  • Traces the eval's filters don't match, and traces outside its sample.
  • Tests in the editor (a judge in a test still calls your provider).
  • A step that couldn't run because of Trodo (the sandbox, or a model it runs, being busy). Those are retried.

When you hit the plan limit

Once the organization has used its Executions for the period, traces aren't scored at all, and are counted under Not scored · over the plan limit until the next period. A walk is never cut off halfway: the decision is made before the first step.

On Pro with Usage beyond your plan turned on (Settings → Billing), and always on Enterprise, evals keep scoring and the extra Executions are billed. See Usage beyond your plan.

Re-scoring history is billed like live traffic, and its estimate is shown before it starts.

Limits

Trace input / output read256 KB each (longer arrives truncated, marked)
Spans loaded per runthe newest 5,000
Runs loaded per conversationup to 200: the first and the newest
Python run time10 seconds per call
What a Python step returns / passes on256 KB each
Judge reasoning kept2,000 characters
Grounded contextthe newest 20 matching spans by default, up to 200
An Intent500 characters
Re-scoring historyup to 50,000 traces per re-score
Audit slice0.5% of traffic by default

States and reasons

Every result says what happened. Passed and failed results reached an ending. The rest:

StateReasonWhat to do
WaitingA Human review stepGrade it in the queue
Erroredpython_exception, python_bad_return, timeout, …See Python steps → When it errors
judge_provider_error, judge_no_model, …See LLM judge → When it errors
cross_encoder_missing_input, cross_encoder_config, …See Semantic checks → When it errors
output_type_mismatch, context_mismatchA step returned something it doesn't declare
unit_missingThe trace was deleted before it was scored

An errored result is never a pass or a fail, and never counts toward a pass rate: evals don't guess.

Traffic that is not scored at all is counted on Analysis under Not scored:

ReasonShown as
filterdid not match the filter
sampledoutside the sample
dedupalready scored
entitlementover the plan limit

On this page