Evals with Lucid
Create evals by describing them, ask why results fail, improve an eval in plain words, and read any eval's results. Lucid works from your real traces.
Lucid knows your evals as well as your traces. It can read them, explain them, build new ones from a sentence, and propose improvements, always as a card you act on rather than a silent change.
Open Lucid from the sidebar, or from Lucid on any eval's page: it opens beside the page already knowing which eval you're looking at.
Create an eval by describing it
Create an eval that catches answers which make up numbers.
Lucid doesn't guess. It:
- Checks whether an eval already does this, then asks whether to use it, change it, or add another.
- Reads your agent: its recent runs, which tools and model calls it makes, what their outputs look like.
- Works out where the problem would show in this agent. For an agent that answers from SQL results, a made-up number is a figure in the answer that isn't in any SQL output.
- Asks only what it can't see: which agent, if several are plausible; which model the judge should use, if you have several providers.
- Designs it as triage (gather the evidence in Python, a Grounded check, a judge only for what that can't settle) and tries it on two real traces.
- Creates it switched off, with a card: its steps, what the tries decided, and Enable.
Name a model to skip the question (judge it with gpt-4.1-nano) or say no judge.
Ask about results
Why did "Recovers from failed tools" fail so much yesterday?
Which evals got worse after the last deploy?
Show me the runs "Went well" failed that people marked right.
Lucid reads results the way the product counts them: pass rate is passed ÷ decided, versions are never pooled, and a run no eval scored isn't a pass. Lists come back as tables that open the eval or the trace.
Improve an eval
Make "Answer quality" stricter about unsupported numbers.
These three should have passed: the user was asking a follow-up, not repeating themselves.
Lucid reads the results in question with every step's answer, decides whether the eval or the agent is wrong, and records a proposed change. The card shows what it changes, what it would flip, Put in production and Reject. When the agent was at fault, the card says what to fix in the agent instead.
From a thumbs-down
Marking results wrong on an eval's page opens Lucid with a message already written: the eval, each result, what it said, what you say it should have been, and your reason. The message asks Lucid to check each one against the trace, say where it disagrees with you, find the step that decided wrongly and propose the fix. Send it as is, or add to it. See Marking results wrong.
What Lucid won't do
- Turn an eval on by itself. Created evals start off; you enable them from the card or Settings.
- Put a change in production. Proposals wait for you.
- Invent results. If something wasn't scored, it says so.
Self-improving evals
An alert on the eval and an agent that has Lucid improve it when the alert fires: proposed changes to the eval, fixes for your agent, and how you decide what ships.
What evals read from your traces
Evals score the traces you already send. The fields that make them sharper (agent names, input and output, conversation ids, span kinds, metadata and feedback) and how to set each from the SDK.