Choosing a step

The five kinds of step an eval is built from: Python, Semantic check, LLM judge, Filter and Human review. What each is for, what it costs, and how it is set up.

Every eval is built from the same five kinds of step. In the editor, add one with + under the step it should follow (or drag one from the rail on the left). Each opens a panel with three tabs:

  • Settings: the step itself, its code, prompt or check.
  • Output & routing: what it answers, and where each answer goes. See Routing and endings.
  • Test: run only this step, or everything up to it, on the trace you picked in Test.
StepAnswersSpeedModel callBest for
PythonWhatever you declare: true/false, a number, one ofMillisecondsNoneExact rules, parsing, counting, gathering evidence for later steps
Semantic checkTrue / falseMilliseconds to secondsNone (small models Trodo runs)Grounded in the evidence? Was it answered? Did the user repeat themselves? Like a known failure?
LLM judgeWhatever you declareSecondsYes, on your provider keyNuance, tone, scope, categories of failure
Filter stepTrue / falseInstantNoneA rule on trace fields, without code
Human reviewOne of the options you listWhen someone grades itNoneWhat the judge was unsure about; building labels

And every walk ends at an ending: a named Pass or Fail.

The same trace, three ways

Take "Did the agent's answer invent a number?"

Built withHowTrade-off
Only a judgeShow the judge the question, the tool output and the answer; ask for a 0 to 1 score.Works, but pays for a judge call on every run, including the thousands whose numbers obviously match.
Only PythonExtract every number from the answer; check each appears in the tool output.Free and exact for figures, blind to made-up claims without numbers.
TriagePython gathers the evidence → Grounded checks every claim, and figures exactly → only unsupported answers reach the judge.Free for most traffic, a judge only where it adds something. This is the Grounded in tool results template.
A larger triage eval: Python finds failed tool calls, Semantic checks test whether the answer rests on the tool output and whether the failure was admitted, judges score the rest, and people review the unsure cases.

On this page