Concepts
The words evals are built from: levels, steps, answers and routes, endings, results, versions and production. Plus how to name things.
Levels: what an eval scores
Every eval scores one level of your traffic, chosen when you create it and fixed after that.
Run
One execution of your agent: the question it got, the answer it gave, and every span inside.
Did it answer the question? Did it invent numbers? Was it too slow?
Span
One step inside a run: a model call, a tool call, a retrieval.
Did this SQL call succeed? Were the tool arguments valid? Was the completion cut off?
Conversation
A whole thread: every run that shares a conversation id, scored again as it grows.
Did the user get what they came for? Did they have to repeat themselves?
An eval reads upward, never sideways. A span eval sees its span, the run it belongs to, and a summary of that run's conversation. A run eval sees the run, all its spans, and its conversation. A conversation eval sees every run in the thread and their spans. What a level cannot see is not offered in the editor, and reading it from Python raises a clear error rather than returning nothing.
When things are scored:
- Runs are scored when the run finishes.
- Spans are scored as each one lands.
- Conversations are scored a couple of minutes after the latest turn, and scored again each time the thread grows: a conversation's verdict is always about the whole thread so far.
Steps, answers and routes
An eval is a small graph. Each step answers one thing, in a declared shape:
| Answer | Example | How routes read it |
|---|---|---|
| True / false | ok = true | True goes to …, False goes to … |
| Number in a range | score = 0.82 (from 0 to 1) | Ranges: below 0.5 → …, 0.5 to 0.8 → …, 0.8 and up → … |
| One of a few options | coverage = partial | One route per option |
A step's routes say where each answer goes: to another step, or to an ending. The first route that matches wins, and every step has an otherwise for anything left. A route can't point back at an earlier step, so every trace's walk ends.
- PythonGather tool results
found = has_evidencehas_evidence → Rests on the tool results - Semantic checkRests on the tool results
ok = false · 3 of 5 claims supportedfalse → Judge the grounding - LLM judgeJudge the grounding
score = 0.31below 0.5 → Invented claims - Fail · Invented claims
HALLUCINATION
The walk a trace takes is its path, shown on every result as step names joined by arrows: Gather tool results → Rests on the tool results → Judge the grounding → Invented claims.
Endings
A walk ends at an ending: Pass or Fail, with a name. Name endings for why: "Grounded (judged)", "Invented claims", "Recovered but did not answer". The name is what shows on results, in trace views and in charts, and a plain "Pass" tells you nothing about which way the trace got there.
A Fail ending also carries a reason code: a short upper-case tag like HALLUCINATION or TOOL_ERROR that results are grouped and searched by. Several endings can share one.
An ending can show a score too: a number from a step on every path to it, usually a judge's, shown beside the verdict as Fail · 0.31 · HALLUCINATION.
Results
Every scored trace gets a result with one of these states:
| State | Meaning |
|---|---|
| Passed / Failed | The walk reached an ending. Only these count toward a pass rate. |
| Waiting | The walk reached a Human review step and is in the grading queue. Grading it finishes the walk. |
| Errored | A step could not run: Python raised, a judge's provider refused, a check had nothing to read. Evals never guess, so an errored result is neither a pass nor a fail, and it says which step failed and why. |
Some traffic is not scored at all, and no result is made: traces the eval's filters don't match, traces outside its sample, and anything over your plan's limit. Analysis counts these under Not scored, with the reason.
Versions and production
Every change to what an eval does is a new version. Saving in the editor asks for a note, what changed, and why, and writes v2, v3, v4. A version is never edited afterwards; restoring an old one writes a new version equal to it.
Saving does not change what runs. One label, production, says which version scores new traffic, and moving it is a separate step: Set as production. So you can save a draft, test it, and ship it when you are ready, and a change Lucid proposes is simply a version that is not in production yet.
v2 is saved. Nothing changes yet: v1 still scores every new run.
Results always record the version that scored them, and versions are never pooled: charts show each version as its own line, and Analysis puts them side by side.
What is not a version: the name, the description, on/off and sampling. Those change in place, immediately.
On and off
An eval that is off scores nothing new; its past results stay. Turn it on or off from Settings in the editor, or from the list.
- An eval created Simple, from scratch is on as soon as it is created.
- An eval created Triage, from scratch stays off until it has a first step.
- An eval created from a template, or by Lucid, starts off, so you can test it first.
Naming conventions
Names are what everyone reads, on results, in trace views, in alerts, in Slack. A few conventions keep them useful:
| Thing | Convention | Good | Avoid |
|---|---|---|---|
| Eval | What passing means, as a claim | Answer rests on tool results · Refund answers stay within policy | Eval 3 · Hallucination check v2 |
| Step | The question it answers | Did it hide the failure? · Rests on the tool results | Step 2 · judge |
| Ending | The outcome and the way there | Grounded (judged) · Recovered but did not answer | Pass on four different endings |
| Reason code | UPPER_SNAKE, one per kind of failure | HALLUCINATION · TOOL_ERROR · GAVE_UP | FAIL · BAD |
| Output field | lower_snake, what the value is | score · coverage · outcome · ok | x · result2 |
| Context key | lower_snake, what it holds | evidence · failed_tools · expected | data · tmp |
| Version note | What changed and why, one line | judge threshold 0.7 → 0.8: too many weak answers passed | update · fix |
Put the agent in a filter, not in the name, unless two evals differ only by agent.
Glossary
| Term | Meaning |
|---|---|
| Eval | A named check on one level of traffic, with a version history. |
| Level | Run, span or conversation: what one result is about. |
| Step | One part of an eval that answers something typed. |
| Route | Where a step's answer sends the walk next. |
| Ending | A named Pass or Fail where the walk stops. |
| Path | The steps one trace walked, ending included. |
| Result / verdict | What an eval decided about one trace. |
| Version | A frozen copy of what an eval does. |
| Production | The one version that scores new traffic. |
| Context | What a Python step passes on to later steps. See Variables. |
| Proposed change | A version (or an agent fix) suggested by Lucid, waiting for a decision. |