Choosing a step
The five kinds of step an eval is built from: Python, Semantic check, LLM judge, Filter and Human review. What each is for, what it costs, and how it is set up.
Every eval is built from the same five kinds of step. In the editor, add one with + under the step it should follow (or drag one from the rail on the left). Each opens a panel with three tabs:
- Settings: the step itself, its code, prompt or check.
- Output & routing: what it answers, and where each answer goes. See Routing and endings.
- Test: run only this step, or everything up to it, on the trace you picked in Test.
| Step | Answers | Speed | Model call | Best for |
|---|---|---|---|---|
| Python | Whatever you declare: true/false, a number, one of | Milliseconds | None | Exact rules, parsing, counting, gathering evidence for later steps |
| Semantic check | True / false | Milliseconds to seconds | None (small models Trodo runs) | Grounded in the evidence? Was it answered? Did the user repeat themselves? Like a known failure? |
| LLM judge | Whatever you declare | Seconds | Yes, on your provider key | Nuance, tone, scope, categories of failure |
| Filter step | True / false | Instant | None | A rule on trace fields, without code |
| Human review | One of the options you list | When someone grades it | None | What the judge was unsure about; building labels |
And every walk ends at an ending: a named Pass or Fail.
The same trace, three ways
Take "Did the agent's answer invent a number?"
| Built with | How | Trade-off |
|---|---|---|
| Only a judge | Show the judge the question, the tool output and the answer; ask for a 0 to 1 score. | Works, but pays for a judge call on every run, including the thousands whose numbers obviously match. |
| Only Python | Extract every number from the answer; check each appears in the tool output. | Free and exact for figures, blind to made-up claims without numbers. |
| Triage | Python gathers the evidence → Grounded checks every claim, and figures exactly → only unsupported answers reach the judge. | Free for most traffic, a judge only where it adds something. This is the Grounded in tool results template. |

Templates
33 ready-made evals for the checks most agents need: answer quality, hallucination, safety, tools, cost and latency, conversations and people in the loop. Plus how to adapt one.
Python steps
Write evaluate(t) over the trace: the fields each level can read, returning an answer and passing context on, the helpers, the libraries, the sandbox and its limits.