Concepts

The words evals are built from: levels, steps, answers and routes, endings, results, versions and production. Plus how to name things.

Levels: what an eval scores

Every eval scores one level of your traffic, chosen when you create it and fixed after that.

Run

One execution of your agent: the question it got, the answer it gave, and every span inside.

Did it answer the question? Did it invent numbers? Was it too slow?

Span

One step inside a run: a model call, a tool call, a retrieval.

Did this SQL call succeed? Were the tool arguments valid? Was the completion cut off?

Conversation

A whole thread: every run that shares a conversation id, scored again as it grows.

Did the user get what they came for? Did they have to repeat themselves?

An eval reads upward, never sideways. A span eval sees its span, the run it belongs to, and a summary of that run's conversation. A run eval sees the run, all its spans, and its conversation. A conversation eval sees every run in the thread and their spans. What a level cannot see is not offered in the editor, and reading it from Python raises a clear error rather than returning nothing.

When things are scored:

  • Runs are scored when the run finishes.
  • Spans are scored as each one lands.
  • Conversations are scored a couple of minutes after the latest turn, and scored again each time the thread grows: a conversation's verdict is always about the whole thread so far.

Steps, answers and routes

An eval is a small graph. Each step answers one thing, in a declared shape:

AnswerExampleHow routes read it
True / falseok = trueTrue goes to …, False goes to …
Number in a rangescore = 0.82 (from 0 to 1)Ranges: below 0.5 → …, 0.5 to 0.8 → …, 0.8 and up → …
One of a few optionscoverage = partialOne route per option

A step's routes say where each answer goes: to another step, or to an ending. The first route that matches wins, and every step has an otherwise for anything left. A route can't point back at an earlier step, so every trace's walk ends.

The walk a trace takes is its path, shown on every result as step names joined by arrows: Gather tool results → Rests on the tool results → Judge the grounding → Invented claims.

Endings

A walk ends at an ending: Pass or Fail, with a name. Name endings for why: "Grounded (judged)", "Invented claims", "Recovered but did not answer". The name is what shows on results, in trace views and in charts, and a plain "Pass" tells you nothing about which way the trace got there.

A Fail ending also carries a reason code: a short upper-case tag like HALLUCINATION or TOOL_ERROR that results are grouped and searched by. Several endings can share one.

An ending can show a score too: a number from a step on every path to it, usually a judge's, shown beside the verdict as Fail · 0.31 · HALLUCINATION.

Results

Every scored trace gets a result with one of these states:

StateMeaning
Passed / FailedThe walk reached an ending. Only these count toward a pass rate.
WaitingThe walk reached a Human review step and is in the grading queue. Grading it finishes the walk.
ErroredA step could not run: Python raised, a judge's provider refused, a check had nothing to read. Evals never guess, so an errored result is neither a pass nor a fail, and it says which step failed and why.

Some traffic is not scored at all, and no result is made: traces the eval's filters don't match, traces outside its sample, and anything over your plan's limit. Analysis counts these under Not scored, with the reason.

Versions and production

Every change to what an eval does is a new version. Saving in the editor asks for a note, what changed, and why, and writes v2, v3, v4. A version is never edited afterwards; restoring an old one writes a new version equal to it.

Saving does not change what runs. One label, production, says which version scores new traffic, and moving it is a separate step: Set as production. So you can save a draft, test it, and ship it when you are ready, and a change Lucid proposes is simply a version that is not in production yet.

Results always record the version that scored them, and versions are never pooled: charts show each version as its own line, and Analysis puts them side by side.

What is not a version: the name, the description, on/off and sampling. Those change in place, immediately.

On and off

An eval that is off scores nothing new; its past results stay. Turn it on or off from Settings in the editor, or from the list.

  • An eval created Simple, from scratch is on as soon as it is created.
  • An eval created Triage, from scratch stays off until it has a first step.
  • An eval created from a template, or by Lucid, starts off, so you can test it first.

Naming conventions

Names are what everyone reads, on results, in trace views, in alerts, in Slack. A few conventions keep them useful:

ThingConventionGoodAvoid
EvalWhat passing means, as a claimAnswer rests on tool results · Refund answers stay within policyEval 3 · Hallucination check v2
StepThe question it answersDid it hide the failure? · Rests on the tool resultsStep 2 · judge
EndingThe outcome and the way thereGrounded (judged) · Recovered but did not answerPass on four different endings
Reason codeUPPER_SNAKE, one per kind of failureHALLUCINATION · TOOL_ERROR · GAVE_UPFAIL · BAD
Output fieldlower_snake, what the value isscore · coverage · outcome · okx · result2
Context keylower_snake, what it holdsevidence · failed_tools · expecteddata · tmp
Version noteWhat changed and why, one linejudge threshold 0.7 → 0.8: too many weak answers passedupdate · fix

Put the agent in a filter, not in the name, unless two evals differ only by agent.

Glossary

TermMeaning
EvalA named check on one level of traffic, with a version history.
LevelRun, span or conversation: what one result is about.
StepOne part of an eval that answers something typed.
RouteWhere a step's answer sends the walk next.
EndingA named Pass or Fail where the walk stops.
PathThe steps one trace walked, ending included.
Result / verdictWhat an eval decided about one trace.
VersionA frozen copy of what an eval does.
ProductionThe one version that scores new traffic.
ContextWhat a Python step passes on to later steps. See Variables.
Proposed changeA version (or an agent fix) suggested by Lucid, waiting for a decision.

On this page