Overview
Failures shows everything that broke, grouped by agent, tool and error. Issues are the problems worth fixing: a root cause a coding agent can work from, the traces that prove it, and a history that tells you when a fix did not hold.
Issues lives under Assess → Issues and has two tabs.
| Tab | What it is | Who makes it |
|---|---|---|
| Failures (the default) | Every hard failure in a period, grouped by agent, tool, error type and category, with charts and every occurrence. | Nobody. It is computed from your traces each time you open it. |
| Issues | The problems worth fixing: one problem in one agent, with a root cause, a suggested change and the runs, spans or conversations that show it. | Lucid writes every issue: after an alert fires, from a Failures row, from an eval, or when you ask it to. |
A failure is a fact about one span or run. An issue is a judgement: these 42 failures are one problem, this is why it happens, and this is what to change. Failures tells you what broke. Issues tells you what to fix.
The loop
- Something failsA tool throws, a run errors, or an eval fails an answer that never errored.
- An alert firesA threshold you set is crossed: 5 failed tool calls in an hour, a p95 over 5 s, 3 looping runs.
- Lucid files itLucid reads what the alert counted, checks what is already filed, and files a new issue or attaches the evidence to an existing one.
- You fix itCopy the issue as a prompt, or hand it to Cursor or Claude Code, which pushes a fix branch.
- Close, and watchWhen the same problem comes back after you close it, the issue reopens as a Regression.
Track issues sets the whole loop up in one click per issue type: the alert with your thresholds, the eval when the failure never errors, and the agent that has Lucid file the issue.
What an issue looks like
An issue from a banking agent, filed by Lucid after a failed-steps alert fired:
| Title | Customer never told after a card freeze: notify_customer fails and is not retried |
| Agent | neobank.fraud_detection |
| Category | Tool |
| Status | Open |
| Root cause | In fraud reviews that end in a freeze, freeze_card succeeds but the follow-up notify_customer call fails ("notify_customer failed") and the agent neither retries nor falls back. The card is frozen with no message to the customer, who finds out at the till and calls support. |
| Suggested change | Retry notify_customer once with backoff after freeze_card. If it still fails, queue the notification and say in the case note that the customer has not been told. |
| Evidence | 20 failing notify_customer spans, each opening its trace |
See Anatomy of an issue for every field, status and category.
Where issues come from
Every issue is written by Lucid, through one store that refuses an issue without evidence or with a root cause too thin to act on. What starts it differs:
| Started by | How |
|---|---|
| An alert | A tracker or any alert wired to a triage agent fires; Lucid investigates exactly the rows the alert counted. |
| The Failures tab | File as issue on a failure row opens Lucid with the failure and its occurrences drafted. |
| An eval | A self-improving eval finds the agent, not the eval, was wrong, and files that finding with the judged traces as evidence. |
| You | Ask it in Lucid: "File an issue for the draft_offer failures on the loan advisor this week." |
See How issues are filed.
In this section
Failures
Anatomy of an issue
How issues are filed
Track issues
Evals, alerts and issues
Fixing an issue
Billing, limits and errors
What evals cost in Executions and in your own provider spend, what is never billed, what happens at the plan limit, the limits of what an eval reads, and every state a result can be in.
Failures
The Failures tab groups every failed span and run by agent, tool, error type and category. What counts as a failure, the ten categories, the charts, the table, and turning a row into an issue or an alert.