Failures
The Failures tab groups every failed span and run by agent, tool, error type and category. What counts as a failure, the ten categories, the charts, the table, and turning a row into an issue or an alert.
The Failures tab is the first thing you see under Assess → Issues. It lists every hard failure in the period, grouped so that order_lookup keeps returning 503 on the checkout agent is one row with a count, not 75 rows. Nothing is stored and no model is involved: it is computed from your traces each time the page loads, and the same failures always group the same way.
A hard failure is one your traces recorded as an error. An answer that is wrong but did not error is a silent failure; evals catch those. See Track issues.
What counts as a failure
Two rules, so each failure is counted once and blamed on the step that broke:
| Counted | Why |
|---|---|
| A failed span with no failed child. | When a tool fails, its parent agent span usually fails too. Counting both would double every failure and blame the agent for what the tool did. Only the deepest failed span counts. |
| A failed run where no span failed. | The run says it failed but no step took the blame. It is listed with Run in the Tool column. |
A span or run is failed when its status is error. What the trace recorded about it (the error type, the message, the HTTP status code) is what the row shows and what decides its category.
How rows are grouped
One row per agent · tool · error type · category. Before grouping, anything that makes one failure look like a thousand is set aside: ids, hashes and numbers in the message. So timed out after 3000ms and timed out after 5000ms land on the same row, while a 503 and a 404 from the same tool stay on different rows because they are in different categories. The most common message names the row.
Categories
Every failure gets one category from what the trace already says. The status code decides first, when the span recorded a real 4xx or 5xx; otherwise the error type and message are read against the rules below, in order, and the first match wins.
| Category | Status code | Or the type or message says | Example |
|---|---|---|---|
| Timeout | 408, 504 | timed out, timeout, ETIMEDOUT, deadline exceeded | TimeoutError: model timed out |
| Rate limited | 429 | rate limit, too many requests, quota exceeded, resource exhausted | 429 Too Many Requests from the model provider |
| Auth | 401, 403 | unauthorized, forbidden, permission denied, invalid api key, credentials | AuthError: invalid api key |
| Connection | ECONNREFUSED, ECONNRESET, ENOTFOUND, socket hang up, network error | connect ECONNREFUSED 10.0.3.12:5432 | |
| Server error | other 5xx | a 5xx in the message, bad gateway, service unavailable, internal server error, upstream | order service returned 503 |
| Not found | 404 | not found, 404, no such | account acct_9f2 not found |
| Refused | refused, guardrail, not allowed, blocked by, policy violation | ToolRefused: table 'payments' is not allowed | |
| Bad request | other 4xx | validation, must be, invalid argument, bad request, schema, missing required | amount must be a positive number |
| Code error | a TypeError, KeyError, ValueError, NullPointerException and the like, or "is not a function", "cannot read properties", "is not defined" | TypeError: Cannot read properties of undefined (reading 'items') | |
| Other | nothing above matched | notify_customer failed |
The tag's colour says what kind of problem it is at a glance: red for broken (Server error, Connection, Code error), amber for slow or throttled (Timeout, Rate limited), blue for rejected (Auth, Refused, Bad request, Not found), grey for Other.
The same categories are an alert filter field, Error category, so an alert on timeouts on this agent counts exactly what this tab counts. Track issues has a template for each.
The period and filters
| Control | |
|---|---|
| Period | The same picker as Traces, defaulting to the last 30 days. Any period up to 90 days; a longer one shows its last 90 days and says so. Up to two days is charted per hour, anything longer per day. |
| Count | The total failures matching the period and filters. |
| Search | Matches the message, span name, tool and error type. |
| Agent, Tool, Category | Multi-selects built from the failures in the period. Selecting any tool leaves run-level failures out, since they have no tool. |
When a period holds more distinct failures than the tab can group, it shows the most frequent ones and says Showing the most frequent failures only. Narrow the period to see the rest.
The dashboard
Four cards above the table count exactly what the table counts, under the same period and filters.
| Card | Shows |
|---|---|
| Over time | Failures per hour or per day, split by Category, Tool or Agent. |
| Breakdown | A donut by Category, Error type or Agent. |
| Tools | A donut of failures by tool. Run-level failures have no tool; the card says how many it leaves out. |
| Failed runs | The share of all runs that failed, per bucket, with the overall figure (4.2% of 18,310 runs). It follows the Agent filter only, since a run has no tool or message to match. |
The table
| Column | |
|---|---|
| Error | The most common message in the group. |
| Category | The category tag, then the error type the trace recorded. |
| Agent | The agent the failures belong to. |
| Tool | The tool or step that failed, or Run for a run-level failure. |
| Count | Failures in the period. Sortable. |
| Users | Distinct users who hit it. A + means at least that many. Sortable. |
| Trend | Failures per day across the period. |
| Last seen | When it last happened. Sortable, and the default sort (newest first). |
Rows load 50 at a time; Load more fetches the next page.
A failure, opened
Click a row to open it in a drawer; the arrows at the top step through the rows.
The header shows the category and error type, the tool (or the agent for a run-level failure), the full message, and Agent, Tool (or Where: Run (no failing span)), Failures, Users, First seen, Last seen and a Per day chart.
Below it, Occurrences lists every failure in the row, newest first: Time, Error, Status, User and Duration. Filter them by message, user or run; Load older fetches more. Click one to open its trace inside the drawer with the failing span selected.
Three actions sit at the top of the drawer:
| Action | What it does |
|---|---|
| File as issue | Opens Lucid with the failure drafted: the agent, tool, category, error type, message, how many and how many users, and up to 20 occurrences by id. Send it and Lucid attaches them to a matching open issue, or reads a few traces and files a new one. See How issues are filed. |
| Create alert | Opens a new alert prefilled for this failure: failed tool calls (or failed runs) on this agent, tool and error type. Set a threshold and a destination and save. See Alerts. |
| Open in Spans / Open in Runs | The same failures in the trace explorer, filtered to errors on this agent and tool, over the same period. |
On a trace
Every run, span and conversation has a Signals tab with two lists: the Issues it is evidence for (a row opens the issue; a merged issue shows as the one it was merged into), and the Failures in it by the rules on this page (a row opens the span). A run also shows the issues its spans are evidence for; a conversation shows its own and its runs'.
Overview
Failures shows everything that broke, grouped by agent, tool and error. Issues are the problems worth fixing: a root cause a coding agent can work from, the traces that prove it, and a history that tells you when a fix did not hold.
Anatomy of an issue
What an issue holds: a title, a root cause and a suggested change written for a coding agent, a category for the kind of fix, a status, the runs, spans and conversations that prove it, and a timeline that keeps every version.