Self-improving evals

An alert on the eval and an agent that has Lucid improve it when the alert fires: proposed changes to the eval, fixes for your agent, and how you decide what ships.

Evals drift. Your agent changes, your users change, and a judge prompt that was right in March passes things in June it shouldn't. A self-improving eval notices and does something about it: when its results start to look wrong, Lucid reads them, decides whether the eval or the agent is at fault, and proposes a fix to the eval, to the agent, or to both. Nothing changes until you approve it.

The eval loopYour agent runs; evals score them; an alert watches what they decide; Lucid investigates; a change is proposed to the eval and to the agent; you decide what ships. Then it goes round again on new traffic.Your agent runsEvals score themAn alert watchesLucid investigatesA change is proposedYou decideon every runnothing ships until you say so

How it works

Self-improving isn't a hidden scheduler. Turning it on creates two ordinary things you can see and edit where they live:

  1. An alert on the evalWhen failures are at or above 5 over 1 hour. It lives in Alerts as Self-improving · your eval.
  2. An agentThe alert starts it; its one step, Improve eval, runs Lucid with your instructions. It lives in Agents as Improve eval · your eval.
  3. Lucid investigatesReads the results behind the alert, what each step answered, and the traces they scored.
  4. A proposed changeA new version of the eval, not in production, and what to fix in the agent. It waits in Proposed changes; with a Slack channel on the alert, it is posted there too.

Turning it on

Open the eval's Settings in the editor and click Deploy self-improving under Self-improving.

The self-improving form: when failures are at or above 5 over 1 hour; what Lucid should do; start paused; Deploy.
Setting
WhenThe measure to watch (Failures, Pass rate, Results marked wrong, Audit disagreements, Errored, Waiting for a person or Average score), how to compare it (at or above, above, below, at or below a value) and over how long (15 minutes to 7 days).
What Lucid should doThe instructions the agent gives Lucid. The default: find why the results came out the way they did; if the eval decided them wrongly, propose the smallest change that decides them rightly, using code where code can settle it and a clearer judge prompt where it can't; if the eval was right and the agent went wrong, change nothing in the eval and say what to fix in the agent.
Start pausedCreate both, switched off.

Deployed, the panel shows Live or Paused, a sentence saying what it watches, links to the Alert and the Agent, and when it last fired.

Eval settings with self-improving paused: Resume, Run now, and 'When failures is at or above 1 over 7 days, Lucid looks at the results and proposes a change', with links to the alert and the agent.
  • Pause / Resume switch the alert and the agent together.
  • Run now runs the agent once, straight away, with the alert's current value, paused or not. The proposal appears in a minute or two.
  • To change the condition, the instructions, or add a Slack step, edit the alert or the agent directly.

What Lucid does

Given the results behind an alert (or the ones you marked wrong, or recent failures):

  1. Reads the evidence: each result, what every step answered, and the trace it scored.
  2. Checks production first. Results scored by an older version are re-checked against the version in production now. If production already decides them the way people said, nothing is proposed, because the fix is already live.
  3. Decides who is wrong. If the eval decided them wrongly, it designs the smallest change that decides them rightly. If the eval was right and the agent misbehaved, it leaves the eval alone.
  4. Tries it. The changed eval is walked on those same traces with its Python run for real, and one failed attempt is retried.
  5. Weighs it. What else the change would flip on recent results, and whether it breaks any result a person labelled.
  6. Records it as a proposed change: a new version, not in production, plus what to fix in the agent whenever the agent was genuinely wrong.

Proposed changes

Every proposal lands in one feed, Evals → Proposed changes: the ones self-improving raised, the ones you asked Ask for, and the ones that came from marking results wrong. New ones are marked; the count beside the button is what's waiting for you.

Proposed changes: a feed by day on the left; on the right a proposal to make the eval more accurate, with Put in production and Reject, why, what it would change, and where it changes the eval.

A proposal to change the eval shows:

  • What kind of change (Make the eval more accurate) and which step it touches: a rule, worked examples for the judge, a new judge prompt, or a new step.
  • Why it is proposing this, in plain words.
  • What it would change to the eval: how many recent results it would flip, each way, and whether it gets any result a person labelled wrong.
  • Where it changes the eval: the step map, with the changed steps marked.
  • Why this came up: the traces that prompted it, with what people said.
Action
Put in productionSets the proposed version as production. If production moved since the proposal was made, the change is re-applied on top of it as a new version. Your re-score policy runs as for any version.
RejectSay why (too broad: it would let real failures through). The reason is kept, and the next proposal learns from it.
RevertFor a change already in production: production goes back to the version before it.

Fixes for your agent

When the eval was right and the agent was wrong, the proposal is Fix the agent: no eval change, and a summary for your engineers of what the agent does wrong, how often, where the eval saw it, and what to change. Not code. The instruction to add, the tool to harden, the error to handle.

A Fix the agent proposal: the SQL generator emits syntax the SQL tool rejects; what to change in the agent's SQL rules.

Mark it Done once your team has fixed it, with a line on what was done (added a rule to the system prompt). A proposal can carry both: a better eval and a fix for the agent.

Staying in control

  • A proposal is a version like any other. Nothing runs until someone puts it in production.
  • Every proposal says what it would flip and whether it contradicts a person's label.
  • Every version records who made it; Lucid's show as authoring. See Who did what.
  • Rejections are remembered. So are your marks on results: the more you label, the better the next proposal.

On this page