Overview
Evals score every run, span and conversation your agent produces. Cheap checks first, a judge or a person only where one is needed, and a proposed fix for what they catch.
An eval answers one question about your agent's work, on real production traffic, every time it happens. Did the answer rest on the tool results, or did it invent a number? Did the user get what they came for? Did that SQL call succeed? Each answer is a verdict: pass or fail, with the reason. You can count them, chart them, alert on them and act on them.
Most eval tools give you one of two things: a rule you write in code, or a model that grades everything. Rules are free but blunt. Grading everything with a model is accurate but expensive, slow, and pays for the easy cases as much as the hard ones. Trodo evals are triage: an eval is a small graph of steps, and each trace walks it until something can decide. A rule settles what a rule can settle. A small, fast model settles most of what is left. Your LLM judge sees only what those could not decide. A person sees only what the judge was unsure about.
Why it matters
Lower cost
A judge call on every trace is the most expensive way to evaluate. Triage sends the judge only what the cheap steps cannot decide, usually a tenth of your traffic or less. That is what makes scoring all of production affordable instead of a sample.
Higher accuracy
Each step does the part it is best at. Code checks what code can check exactly: a status, a field, a number. Semantic checks read meaning: is this claim in the tool output? A judge handles nuance. A person settles the genuinely unclear cases, and their answers become labels.
It gets better on its own
When an eval starts deciding badly (too many failures, or results people keep marking wrong) Lucid reads the evidence and proposes a better version, with what it would change and whether it breaks anything a person already labelled. Nothing goes live until you say so.
It improves your agent
When the eval was right and the agent was wrong, the proposal says so, and writes the fix your engineers should make to the agent: the instruction to add, the tool to harden, the retry to handle. Evals stop being a dashboard and become a to-do list.
The loop
Every part of this is something you can open, read and change. Evals score your traffic; an alert watches what they decide; when it fires, Lucid investigates and proposes a change to the eval and to your agent; you decide what ships.
What an eval looks like
An eval scores one level of your traffic: a run, a span, or a whole conversation. It is built from steps. Each step answers something typed (true or false, a score, one of a few categories), and its answer picks where the trace goes next, until it reaches an ending: a named pass or fail.

Watch one trace walk it:
- PythonGather tool results
found = has_evidencehas_evidence → Rests on the tool results - Semantic checkRests on the tool results
ok = false · 3 of 5 claims supportedfalse → Judge the grounding - LLM judgeJudge the grounding
score = 0.31below 0.5 → Invented claims - Fail · Invented claims
HALLUCINATION
There are five kinds of step. Use one on its own (a Simple eval is one step, plus a rule for which answers pass) or chain several into a Triage eval.
| Step | What it is | Costs a model call? |
|---|---|---|
| Python | Your own evaluate(t) over the trace. Exact, instant, anything code can check. | No |
| Semantic check | Small models that read meaning: Grounded (is every claim supported?), Intent (was it answered? did the user repeat themselves?), Matches (is it like a known example?). | No |
| LLM judge | A model you choose, on your own provider key, answering your prompt. | Yes, on your key |
| Filter | A rule on trace fields, no code. | No |
| Human review | The trace waits in the grading queue for someone on your team. | No, a person's time |