Results and analysis

Read what an eval decided: every verdict and its path, Analysis with versions side by side, marking results wrong, the scores in your trace views, searching traces by eval results, and alerts on evals.

An eval opens on its results. The header says which version is in production and whether the eval is on, and holds Lucid, Create alert, Versions and Edit. Below it, three tabs share one set of controls:

Control
Verdicts · Queue · AnalysisEvery result; this eval's items waiting for a person; the numbers and charts.
FilterNarrow all three tabs to traces whose walk went through chosen Steps, or took a chosen Path.
VersionAll versions, or one version's results only.
WindowLast 24 hours, 7, 30 or 90 days.

Verdicts

Verdicts: each row is a run with its verdict, the version that scored it, the path it took, when, and thumbs.

One row per result, newest first: the run, span or conversation it scored; the verdict (Pass · 0.889, Fail · 0.667 · TOOLS_FAILED) with the version that decided it; the path; when; and Correct?. Filter by All · Passed · Failed · Errored; the count on the right is what is listed.

Click a row to open its path:

A result's path: each step with what it answered and how long it took, the Semantic check's reading and the text it read, and the ending with its reason code.

Each step shows what it answered, its time and cost, the judge's reasoning, a Semantic check's reading (8 of 15 claims supported, read as: answered · 0.00) and what it passed on. The end shows the ending's name and reason. Open run shows the full trace.

Marking results wrong

An eval is only as good as its verdicts, and you are the best judge of them. On any decided result, 👍 says the eval got it right; 👎 opens What should this eval have said?: pick Pass or Fail and say why.

The thumbs-down form: what the eval should have said, and why it was wrong.

To mark several at once, tick their boxes (or the one in the header, for the whole page). A bar appears at the bottom: should have been The other way, Pass or Fail, one reason for all, Mark N wrong.

Two results selected, with the bar at the bottom: should have been the other way, pass or fail; the reason; Mark 2 wrong.

Either way, Lucid opens beside the page with a message ready to send: the eval, each result with what the eval said and what you say it should have been, your reason, and the instructions: read each result and its trace, say where it disagrees with you, find the step that decided wrongly, and propose the change. Edit it or send it. The proposal lands in Proposed changes.

Every mark is also a label: it counts toward Marked wrong on Analysis, toward the Results marked wrong alert measure, and as evidence the next time the eval is improved.

Analysis

The same window as numbers and charts.

Analysis: summary tiles; versions side by side; pass rate by version over time; verdicts over time; where runs end; cost; step latency; a step's answers and score spread; every path; every step.
CardWhat it shows
TilesPass rate (passed ÷ decided), verdicts, failed, errored and waiting, provider spend and units, and Not scored with the main reason.
VersionsEvery version side by side; see below.
Pass ratePer version, each its own line; a dashed line marks where a version went into production. Pass rate counts decided results only: waiting and errored are states, not failures.
VerdictsPassed, failed, errored and waiting over time.
Where runs endThe step whose decision ended each run, and the ending it led to. A run that ends at a cheap step never reached the judge. That is triage at work.
CostProvider spend (your judge calls) or Units (steps run), by live traffic, re-scoring and audit. Two bills, never added.
Step latencyCalls, p50, p95, p99 per kind of step.
PathsEvery walk the eval took, its share, how those runs ended. Click one to see just those runs.
StepsEvery step: runs, most common answer, errors, p50, spend. Click one to add a card for it: its answers and how the runs that got each one ended, its routes, its answers over time, and, for a score, how the scores spread.
Screen accuracyFrom the audit slice: how often each screen's shortcut disagreed with the full path.

Charts group by hour, 6 hours, 12 hours or day on your clock; each has its own Cluster by.

Versions side by side

Versions side by side: pass rate per version; a table of results, pass rate, errored, marked wrong and agreement, cost per verdict and latency; and on the same traces, v1 to v2: 23 traces, agree on 60.9%, 9 fail to pass.

Every version over the same window: results, pass rate, errors, how many results people marked wrong and how often they agreed with it, cost per verdict, and latency. Click a row to see the whole page for that version alone: every number is then that version's, compared with production: Pass rate · v1 47.8% · −39.1 pts vs v2.

Two versions live at different times score different traffic, so their pass rates aren't a like-for-like comparison. On the same traces is: for every pair of versions that scored the same traces (after a re-score), how many they both decided, how often they agree, and which way the rest flipped: v1 → v2 · 23 traces · agree on 60.9% · 9 fail → pass.

In your trace views

Open any run, span or conversation and go to Scores. The Evals section lists every eval that scored it, with the ending's name and what it decided, and 👍 / 👎.

The Evals section of a run's Scores tab: five evals with their endings, Pass, Poor answer 0.42, Recovered but did not answer, Acceptable SQL 0.8, and thumbs.

Expand a row to see each step, what it answered and what it measured. A run lists its spans' results under Spans in this run; a conversation lists its runs' under Runs in this conversation. The Scores badge counts all of them.

Searching traces by eval results

The search on Traces understands eval results (last 90 days):

SearchFinds
"Went well"=failRuns the Went well eval failed
"Answer quality">0.7Runs that eval scored above 0.7
HALLUCINATIONAny result with that reason code, ending name or step name
ok=falseAny step that answered false
similarity>0.9, score<0.5What a step measured
failed, passed, erroredBy outcome

Quote a name that has spaces in it. Combine terms with spaces, commas or and: "Went well"=fail and "Tool success">0.8.

Alerts on evals

Create alert on an eval's page opens a new alert already scoped to it. Watch:

Measure
FailuresFailed results in the window
Pass ratePassed ÷ decided
DecisionsDecided results
Errored · Waiting for a person
ScoreAverage, min, max or a percentile of the ending's score
Marked wrongResults people marked wrong
Audit disagreementsWhere a screen and the full path disagreed

Re-scored history never counts toward an alert. An alert on an eval can also start an agent: that is exactly what self-improving sets up. See Alerts.

On this page