Your first eval

From a template to verdicts on live traffic: pick, adapt, test on real traces, turn on, read the results.

This walks through one eval end to end: Answers the question, a judge that scores whether each run's answer addresses what the user asked. It takes a few minutes. Everything after it (Triage, Semantic checks, self-improving) builds on the same five moves.

You need an agent that is already sending traces, and a model provider added in Settings โ†’ Models for the judge to run on.

Open New eval

Go to Evals and click New eval. A pop-up opens with every template, grouped by what they check, and a search at the top: type hallucination, tool errors, frustration, JSON.

The New eval pop-up: a search box, category chips, template cards with their level and the kinds of step they use, and Start from scratch and Create with Lucid buttons.

Pick Answers the question. You could also Start from scratch, or Create with Lucid and describe what you want in your own words. See Evals with Lucid.

Adapt it

The create page opens filled in from the template: its name, what it scores (a Run), and its one step. The box at the top says when to use it and what to change for your agent.

The create page filled in from the Answers the question template: the template's advice, the name, the level, filters, sampling, and the judge's model and prompt.

This template is Simple: one step, plus a rule for which answers pass. You edit it right here:

  • Filters: leave empty to score every agent, or add Agent name is support-bot to score one.
  • Model: the judge runs on your provider key. Trodo picks the provider you used last; change it if you like.
  • Judge prompt: {{run.input}} and {{run.output}} are filled in from each run. Click a field in the list below the prompt to insert another.
  • Decision: Passes when the score is from 0.7 to 1. Raise it if your answers must be exact.
The Simple form: eval type, model, judge prompt, the field list, the output (a score from 0 to 1), the decision (passes from 0.7 to 1) and the fail reason.

Click Create and open. An eval made from a template starts switched off: nothing is scored until you turn it on.

Test it on real traces

In the editor, click Test. Pick a recent run on the left and Run the whole eval. You see every step, what it answered, where it went, and the verdict, on a real trace, before the eval scores anything.

A test run: each step with its answer, time and cost, the judge's reasoning, and the verdict.

A test writes nothing and is never billed as an eval, but a judge step really calls your provider, so it costs what one call costs. Try a few runs you know are good and a few you know are bad.

Turn it on

Open Settings in the editor and set Enabled to On. From now on every new run that matches the filters is scored. Past runs are not, unless you re-score history.

Read the results

Go back to the eval (the breadcrumb, or Evals โ†’ the eval). An eval opens on its results: every verdict, newest first, with the path it took. Open one to see each step's answer and the judge's reasoning.

The results page: tabs for Verdicts, Queue and Analysis; a filter, a version picker and a time window; a row per result with its verdict, version, path and thumbs.

If the eval got one wrong, click ๐Ÿ‘Ž on it, say what it should have been and why. Lucid opens beside the page with that result ready to investigate. See Marking results wrong.

Next

On this page