Track issues
Every issue type you can track: hard failures at span and run level, one per error category, latency, and silent failures caught by evals. What one click deploys, the thresholds, and how to pause, edit or remove a tracker.
Track issues turns an issue type into a running pipeline in one click. Pick a type, say how much is too much, and Trodo deploys everything that type needs: an alert with your thresholds, an eval when the failure never errors, and an agent that has Lucid file or attach the issue when the alert fires.
Open it from the Issues tab: Track issues, top right.
What one click deploys
- Eval (silent failures only)Catches what never errors, on every matching trace: a tool that answers empty, a loop, a user who had to repeat themselves.
- AlertYour threshold: 5 failed tool calls within 1 hour. For a silent failure, it counts the eval's failures.
- Triage agentStarts when the alert fires. Its step is Lucid, investigating exactly what the alert counted.
- IssueLucid files a new issue, or attaches to the one already open. Optionally posted to Slack.
Each piece is an ordinary object, named so it reads correctly where it lives, and edited there:
| Deployed | Where | Named |
|---|---|---|
| The alert | Alerts | Track · Tool call failures · neobank.fraud_detection |
| The agent | Agents | Track · Tool call failures · neobank.fraud_detection |
| The eval (silent failures) | Evals | Tools answer empty · neobank.fraud_detection |
Nothing about a tracker is stored anywhere else: the Tracking tab reads each one back from its alert, agent and eval. Change the threshold on the alert and the tracker shows the new one.
Deploying one
Click a type to open its settings under it.
| Setting | |
|---|---|
| Agent | All agents, or one. One tracker per agent gives each agent its own alert and its own issues. |
| Fire at | The threshold, in the type's own unit: 5 failed tool calls, 10 % of runs failing, 5000 ms (p95). |
| Within | The alert's window: 5 min, 15 min, 30 min, 1 hour (the default), 6 hours or 24 hours. |
| Only with at least | For rates and percentiles only: the fewest runs or spans in the window before the alert may fire, so 1 failure out of 2 runs is not a 50% error rate. |
| Post to Slack | Optional. The agent posts what Lucid did: the issue it filed or added to, and why. |
Deploy creates everything and switches to the Tracking tab. Deploying the same type on the same agent again resumes the tracker you already have rather than making a second.
Hard failures
Failures your traces recorded as errors, counted the way the Failures tab counts them. Each type is an alert on failed spans or runs; the default is to fire above 5 within an hour.
At span level
One step failing. Failed steps catches all of them; the rest narrow it to one kind of step or one error category, so an issue about rate limits is never mixed with one about bad requests.
| Type | Counts | Default |
|---|---|---|
| Failed steps | Any step that errors: tools, model calls, retrieval | 5 failed steps |
| Tool call failures | Tool calls that error | 5 failed tool calls |
| Model call failures | Calls to the model provider that error | 5 failed model calls |
| Retrieval failures | Retrieval and search steps that error | 5 failed retrievals |
| Timeouts | Steps that time out | 5 timeouts |
| Rate limits | Steps refused with a rate limit (429) | 5 rate-limited steps |
| Auth errors | Steps refused for credentials or permissions (401, 403) | 5 auth errors |
| Server errors | Steps where a service returned a 5xx | 5 server errors |
| Bad requests | Steps a service rejected as invalid (4xx) | 5 bad requests |
| Not found | Steps that asked for something that is not there (404) | 5 not-found errors |
| Connection errors | Steps that could not reach a service | 5 connection errors |
| Refused | Steps a model, guardrail or service refused | 5 refusals |
| Code errors | Steps that threw in the agent's own code | 5 code errors |
At run level
The whole run. Useful when failures do not surface on a step, or when you care about the outcome rather than which step broke.
| Type | Counts | Default |
|---|---|---|
| Failed runs | Runs that end in an error | 5 failed runs |
| Runs with a failed step | Runs where any step failed, even when the run recovered | 5 runs |
| Error rate | The share of runs that fail | 10% of runs failing, with at least 20 runs |
Latency
Slowness is not an error, but it is a problem your users feel. Both fire on the 95th percentile, so one slow outlier does not fire them.
| Type | Counts | Default |
|---|---|---|
| Slow runs | p95 run duration | 30000 ms, with at least 10 runs |
| Slow tool calls | p95 duration of tool calls | 5000 ms, with at least 10 tool calls |
Silent failures
Nothing errored, and the agent still let the user down. These need an eval to see them, so each type deploys one from the eval templates, an alert on its failures (at or above the threshold), and the triage agent. The evals use Python or Semantic checks and never an LLM judge, so there is no model to choose.
| Type | The eval fails | Checks | Default |
|---|---|---|---|
| Tools answer empty | Tool calls that report success but return nothing: an empty output, or one that says "no results" | Span | 5 empty tool results |
| Retrieval finds nothing | A retrieval that returns no results while the agent answers anyway, the classic source of made-up answers | Span | 5 empty retrievals |
| Tool loops | Runs that call the same tool with the same input again and again | Run | 3 looping runs |
| Retries run out | Runs where one tool failed three or more times in a row and never succeeded after | Run | 3 runs |
| Broken model calls | Model calls that error, return no text, or are cut off at the token limit | Span | 5 model calls |
| Users repeat themselves | Conversations where the user asked the same thing again | Conversation | 5 conversations |
| Ends without an answer | Conversations that end on a question the agent never answered | Conversation | 5 conversations |
| Users leave after a failure | Conversations whose last turn failed or refused, and the user left | Conversation | 5 conversations |
A silent-failure issue's evidence is the traces the eval failed, and each shows the eval's verdict and reason in the issue. See Evals, alerts and issues.
The Tracking tab
Every tracker on the team, with its threshold and window, when it last fired, and links to its alert, eval and agent:
Tool call failures · 5 failed tool calls within 1 hour · last fired 2h ago · alert · agent
The switch pauses or resumes a tracker: its alert, its agent and its eval together. A paused eval scores nothing, so a paused tracker spends nothing.
| To | Do |
|---|---|
| Change the threshold or window | Edit the alert on Alerts. |
| Change what Lucid is told, or add a Slack or GitHub step | Edit the agent on Agents. |
| Change what the eval checks | Edit the eval on Evals. |
| Stop tracking | Switch it off here, or delete the alert, agent and eval where they live. |
When the period's Executions are used up, the modal says so: alerts still fire, but the agents and evals do not run until the next period or an upgrade.
What to track first
- Failed steps and Failed runs on each agent you care about. Between them they catch every hard failure, and Lucid splits them into issues by cause.
- A category type (Rate limits, Auth errors, Timeouts) where one kind of failure needs its own threshold or its own Slack channel: rate limits on a payments agent matter at 1, not 5.
- Error rate for busy agents, where a fixed count means something different at 10am and at 3am.
- A silent-failure type for each way your agent can fail without erroring. Tools answer empty and Users repeat themselves catch the most.
How issues are filed
Lucid writes every issue: after an alert fires, from a Failures row, from an eval's finding, or when you ask. How it finds the evidence, avoids duplicates, clusters by cause, and what it writes for the coding agent.
Evals, alerts and issues
Evals decide what is wrong, alerts decide how much is too much, and Lucid decides what the problem is. How an eval result becomes evidence on an issue, how any alert can file issues, and the agent templates that wire them together.