Evals, alerts and issues
Evals decide what is wrong, alerts decide how much is too much, and Lucid decides what the problem is. How an eval result becomes evidence on an issue, how any alert can file issues, and the agent templates that wire them together.
Three parts, each doing one job:
| Part | Decides | On its own it gives you |
|---|---|---|
| Eval | Whether one run, span or conversation was right, including when nothing errored. | A verdict per trace, and a pass rate. |
| Alert | When there is enough of something to act on: your threshold, over your window. | A notification. |
| Issue | What the problem is, why, and what to change. Written by Lucid. | Something a person or a coding agent can fix. |
Issues have no thresholds of their own. The alert is the threshold: an issue is filed when an alert fires and an agent hands it to Lucid. That is what Track issues sets up, and you can wire any alert the same way.
Any alert can file issues
An alert on its own only notifies. To have it file issues, give it an agent whose step is Lucid with the task Investigate an alert into issues. Two ways:
- Track issues deploys the alert and the agent together, for the common types.
- Agents → From template → Triage alerts into issues builds the agent for alerts you already have. Choose Alerts that start it (one agent can watch several), optionally a Slack channel, and Create agent. It is created as a draft: review it and publish.
The Lucid step has two settings: Rows to read at most (40 by default), how much of what the alert counted Lucid reads before filing, and optional Instructions, such as Always file timeouts on the payments tool as infrastructure. What Lucid does with it is on How issues are filed.
The step's output carries the issue (its id, and whether it was created, attached or reopened), so later steps can use it. The other templates in the gallery do:
| Template | After Lucid files |
|---|---|
| Triage alerts into issues | Optionally posts to Slack. |
| Triage, then open a GitHub issue | For each new issue, opens one in your repository with the root cause and suggested change. |
| Triage, then hand to your coding agent | Hands each new issue to Cursor or Claude Code, which pushes a fix branch for you to review. See Fixing an issue. |
| Weekly issue digest | Not an alert: every Monday Lucid summarises what is open, what regressed and what is new, and posts it to Slack. |
From eval results to an issue
An eval catches what an error never will. There are three ways its failures become an issue.
A silent-failure tracker
The simplest. Track issues → Silent failures deploys an eval (Tools answer empty, Users repeat themselves…), an alert on its Failures, and the triage agent. When the alert fires, Lucid reads the traces the eval failed and files or attaches, with each piece of evidence noted with the eval and its reason: Tools answer empty: EMPTY_RESULT.
An alert on any eval
For an eval you built yourself, Create alert on the eval's page opens an alert already scoped to it; watch Failures at or above a count. Then start the Triage alerts into issues agent from it, as above. Lucid treats the eval's failed traces exactly as it treats failed spans. See Alerts on evals.
A self-improving eval
A self-improving eval has Lucid investigate when its results look wrong. Sometimes the eval is right and the agent is wrong: it quoted a rate before the credit profile arrived, it confirmed a transfer to the wrong bank. The proposal then carries a fix for the agent, and that finding is also filed as an issue:
| The issue's | Comes from |
|---|---|
| Title and root cause | The finding: what is wrong in the agent, and why. |
| Suggested change | The change to make, when Lucid knows one. |
| Category | The kind of fix the finding names: prompt, tool, data, code… |
| Evidence | The traces the eval judged, each noted with the eval's reason. |
| Timeline | Filed, with From an eval linking the eval. |
It is filed per agent, and only for an agent with enough of the evidence to reach the eval's threshold, so one odd trace never becomes an issue. When the same finding comes up again, its evidence is added to the same issue; if that issue was closed, it reopens as a Regression.
Example: the Loan advice triage eval keeps failing runs of neobank.loan_advisor. Self-improving finds the eval is right: the system prompt allows an "indicative" rate from stated income alone. It files Rate quoted without pulling the live credit profile, category Prompt, with the 30 failed runs as evidence and the suggested change state a rate only after fetch_credit_profile succeeds.
Eval results on every issue
Whatever filed an issue, each piece of its evidence shows the evals that failed on that trace, beside the error. An issue filed from a hard failure can show that the same runs also failed Grounded in tool results; click the chip to open that eval result. A coding agent reading the issue over MCP or from Copy as prompt sees the same.
Putting it together
A team that tracks both kinds of failure on its loan agent ends up with:
- Hard failuresTrack · Failed steps · neobank.loan_advisor: above 5 failed steps within 1 hour, then Lucid files.
- Silent failuresTrack · Tools answer empty · neobank.loan_advisor: the eval, an alert at 5, then Lucid files.
- QualityLoan advice triage, self-improving: a better eval when the eval is wrong, an issue when the agent is.
- ReviewWeekly issue digest to #loans-eng every Monday.
All of them file into the same Issues tab. Before filing from an alert, Lucid checks what is already filed, including issues an eval filed, so a problem both saw gets the alert's evidence on the existing issue. Any duplicates that slip through, ask Lucid to merge: review the open issues on the loan advisor and merge any that are the same problem.
Track issues
Every issue type you can track: hard failures at span and run level, one per error category, latency, and silent failures caught by evals. What one click deploys, the thresholds, and how to pause, edit or remove a tracker.
Fixing an issue
Copy an issue as a prompt for any coding agent, hand it to Cursor or Claude Code to push a fix branch, read it over MCP, then close it, and learn when the fix did not hold.