Issues

Track issues

Every issue type you can track: hard failures at span and run level, one per error category, latency, and silent failures caught by evals. What one click deploys, the thresholds, and how to pause, edit or remove a tracker.

Track issues turns an issue type into a running pipeline in one click. Pick a type, say how much is too much, and Trodo deploys everything that type needs: an alert with your thresholds, an eval when the failure never errors, and an agent that has Lucid file or attach the issue when the alert fires.

Open it from the Issues tab: Track issues, top right.

What one click deploys

  1. Eval (silent failures only)Catches what never errors, on every matching trace: a tool that answers empty, a loop, a user who had to repeat themselves.
  2. AlertYour threshold: 5 failed tool calls within 1 hour. For a silent failure, it counts the eval's failures.
  3. Triage agentStarts when the alert fires. Its step is Lucid, investigating exactly what the alert counted.
  4. IssueLucid files a new issue, or attaches to the one already open. Optionally posted to Slack.

Each piece is an ordinary object, named so it reads correctly where it lives, and edited there:

DeployedWhereNamed
The alertAlertsTrack · Tool call failures · neobank.fraud_detection
The agentAgentsTrack · Tool call failures · neobank.fraud_detection
The eval (silent failures)EvalsTools answer empty · neobank.fraud_detection

Nothing about a tracker is stored anywhere else: the Tracking tab reads each one back from its alert, agent and eval. Change the threshold on the alert and the tracker shows the new one.

Deploying one

Click a type to open its settings under it.

Setting
AgentAll agents, or one. One tracker per agent gives each agent its own alert and its own issues.
Fire atThe threshold, in the type's own unit: 5 failed tool calls, 10 % of runs failing, 5000 ms (p95).
WithinThe alert's window: 5 min, 15 min, 30 min, 1 hour (the default), 6 hours or 24 hours.
Only with at leastFor rates and percentiles only: the fewest runs or spans in the window before the alert may fire, so 1 failure out of 2 runs is not a 50% error rate.
Post to SlackOptional. The agent posts what Lucid did: the issue it filed or added to, and why.

Deploy creates everything and switches to the Tracking tab. Deploying the same type on the same agent again resumes the tracker you already have rather than making a second.

Hard failures

Failures your traces recorded as errors, counted the way the Failures tab counts them. Each type is an alert on failed spans or runs; the default is to fire above 5 within an hour.

At span level

One step failing. Failed steps catches all of them; the rest narrow it to one kind of step or one error category, so an issue about rate limits is never mixed with one about bad requests.

TypeCountsDefault
Failed stepsAny step that errors: tools, model calls, retrieval5 failed steps
Tool call failuresTool calls that error5 failed tool calls
Model call failuresCalls to the model provider that error5 failed model calls
Retrieval failuresRetrieval and search steps that error5 failed retrievals
TimeoutsSteps that time out5 timeouts
Rate limitsSteps refused with a rate limit (429)5 rate-limited steps
Auth errorsSteps refused for credentials or permissions (401, 403)5 auth errors
Server errorsSteps where a service returned a 5xx5 server errors
Bad requestsSteps a service rejected as invalid (4xx)5 bad requests
Not foundSteps that asked for something that is not there (404)5 not-found errors
Connection errorsSteps that could not reach a service5 connection errors
RefusedSteps a model, guardrail or service refused5 refusals
Code errorsSteps that threw in the agent's own code5 code errors

At run level

The whole run. Useful when failures do not surface on a step, or when you care about the outcome rather than which step broke.

TypeCountsDefault
Failed runsRuns that end in an error5 failed runs
Runs with a failed stepRuns where any step failed, even when the run recovered5 runs
Error rateThe share of runs that fail10% of runs failing, with at least 20 runs

Latency

Slowness is not an error, but it is a problem your users feel. Both fire on the 95th percentile, so one slow outlier does not fire them.

TypeCountsDefault
Slow runsp95 run duration30000 ms, with at least 10 runs
Slow tool callsp95 duration of tool calls5000 ms, with at least 10 tool calls

Silent failures

Nothing errored, and the agent still let the user down. These need an eval to see them, so each type deploys one from the eval templates, an alert on its failures (at or above the threshold), and the triage agent. The evals use Python or Semantic checks and never an LLM judge, so there is no model to choose.

TypeThe eval failsChecksDefault
Tools answer emptyTool calls that report success but return nothing: an empty output, or one that says "no results"Span5 empty tool results
Retrieval finds nothingA retrieval that returns no results while the agent answers anyway, the classic source of made-up answersSpan5 empty retrievals
Tool loopsRuns that call the same tool with the same input again and againRun3 looping runs
Retries run outRuns where one tool failed three or more times in a row and never succeeded afterRun3 runs
Broken model callsModel calls that error, return no text, or are cut off at the token limitSpan5 model calls
Users repeat themselvesConversations where the user asked the same thing againConversation5 conversations
Ends without an answerConversations that end on a question the agent never answeredConversation5 conversations
Users leave after a failureConversations whose last turn failed or refused, and the user leftConversation5 conversations

A silent-failure issue's evidence is the traces the eval failed, and each shows the eval's verdict and reason in the issue. See Evals, alerts and issues.

The Tracking tab

Every tracker on the team, with its threshold and window, when it last fired, and links to its alert, eval and agent:

Tool call failures · 5 failed tool calls within 1 hour · last fired 2h ago · alert · agent

The switch pauses or resumes a tracker: its alert, its agent and its eval together. A paused eval scores nothing, so a paused tracker spends nothing.

ToDo
Change the threshold or windowEdit the alert on Alerts.
Change what Lucid is told, or add a Slack or GitHub stepEdit the agent on Agents.
Change what the eval checksEdit the eval on Evals.
Stop trackingSwitch it off here, or delete the alert, agent and eval where they live.

When the period's Executions are used up, the modal says so: alerts still fire, but the agents and evals do not run until the next period or an upgrade.

What to track first

  • Failed steps and Failed runs on each agent you care about. Between them they catch every hard failure, and Lucid splits them into issues by cause.
  • A category type (Rate limits, Auth errors, Timeouts) where one kind of failure needs its own threshold or its own Slack channel: rate limits on a payments agent matter at 1, not 5.
  • Error rate for busy agents, where a fixed count means something different at 10am and at 3am.
  • A silent-failure type for each way your agent can fail without erroring. Tools answer empty and Users repeat themselves catch the most.

On this page