Overview

Trodo is the observability layer that makes agents better: evals on all of your production traffic, cheap checks first, a judge only where one is needed, and the fix for what they catch.

An agent only gets better if something is measuring it. Trodo scores every run in production, using the cheapest thing that can decide it, and turns what fails into a fix for your agent.

How Trodo makes an agent betterEvery run in production is scored. A rule settles most of it for nothing, a small model settles most of what is left, your judge sees only the hard cases and a person only what the judge was unsure about. What fails becomes a fix to the agent, and the number of failing runs falls week over week.EVERY RUNSETTLED BY THE CHEAPEST STEP THAT CAN DECIDE ITA rulefreeA small modelfreeYour judge$A person1 min✓ passed✕ failedFAILING RUNS, BY WEEKA fix for your agent

The failures that cost you never throw

Agents rarely crash. They make up a figure, answer a question nobody asked, drop the constraint from three turns ago, answer from a tool that quietly returned an error, promise something you cannot honour, or send the user round in circles until they give up. No stack trace, latency chart or error rate sees any of it. Only something that reads what the agent said and did: an eval, on real traffic rather than a test set you froze months ago.

Why almost nobody scores production

A judge on every run doubles your model bill to be told 19 runs in 20 were fine, and it drifts without telling you. Code is exact and free, but it cannot tell you whether an answer was supported. So teams score a sample, and the sample decides which failures they get to hear about.

Scoring a tenth of your traffic does not make a failure a tenth as common. It makes you a tenth as likely to see it.

failed failed and scored scored, fine never looked at
Failures in 300 runs5
You see00% of them
People it hits before you see one~10at 100% coverage: 1

Sampling never reduces failures. It reduces sightings, and delays the ones it leaves you. Drag the coverage down and watch the red squares nobody ever reads.

Triage: the expensive step runs only when it is needed

Most eval tools are flat: every evaluator runs on every trace, and you compose the results at the end, so the judge grades the empty answers and the "yes, thanks" replies that a two-line rule had already settled.

A Trodo eval is a small graph of steps, and each trace walks it until something can decide. Same verdicts, a fraction of the bill, which is what makes scoring all of production affordable.

Fixes for your agent, not one more dashboard

When an eval catches something, Trodo reads the results and proposes the change: a better version of the eval when the eval was wrong, and when the agent was wrong, what to fix in the agent, with the runs that prove it. Repeat failures move down the ladder into a rule, so the judge stops paying to rediscover them.

Detection that ends in a chart changes nothing. Detection that ends in a change to your agent is how it gets better.


Everything the loop is built on

Observability

Everything starts with the trace. One wrap around your agent captures every run as a tree of spans: each LLM call, tool call, and nested step, with inputs, outputs, tokens, cost, latency, and errors. Works with any stack (OpenAI, Anthropic, LangChain, Vercel AI SDK, raw HTTP) through the Trodo SDK or OpenTelemetry.

  • One wrap, full trace. wrapAgent records the whole run as nested spans; auto-instrumentation captures provider calls with no extra wiring.
  • Runs, conversations and users. Bind runs to a user and a thread, stitch multi-step sessions across processes and services.
  • The data layer for everything else. Evals, Issues, Capabilities and Lucid all read these same traces. There is nothing extra to send for any of them.
Wrap your agent once and capture every run as a tree of spans: calls, tools, tokens, cost, and errors.

Evals

Score runs, spans and whole conversations on real traffic. Five kinds of step (Python, semantic checks, an LLM judge on your own key, a filter and human review) compose into one graph that ends at a named pass or fail you can chart, alert on and act on.

  • 33 templates for the usual suspects: hallucination, tool failures, safety, frustration, cost.
  • Versions: test on real traces, save a version, put it in production, re-score your history with it.
  • Results: pass rate over time, cost per verdict, versions side by side, and a thumbs-down that opens Lucid with the whole story.
Score every run, span and conversation: cheap checks first, a judge only where one is needed.

Lucid

Most observability tools make you the query engine: filter, scroll, read, repeat. Lucid answers in plain language instead, across your traces and your evals. It builds evals from a sentence, explains why results failed, and proposes improvements, always as a card you act on rather than a silent change.

  • Natural-language investigation across runs, spans, events, evals and users.
  • Scoped to anything: a user, a thread, a cohort, a single run, one eval.
  • It writes evals too. Describe what to catch; Lucid reads your agent, designs the triage, tries it on real traces, and creates it switched off.
Query traces, runs, journeys and eval results in plain language. No filters, no SQL.

Issues

Evals answer the questions you chose to ask. Failures and Issues cover what breaks without one. The Failures tab, where Issues opens, groups every hard failure in your traces (a tool that threw, a model call that timed out, a rate limit) by agent, tool and error category, with charts of when they happen. An issue is one problem worth fixing: Lucid files it from the failing traces with a root cause, a category (code, prompt, tool, infrastructure, data or model) and the evidence, and keeps it current as more arrives.

  • Track issues in one click: a template sets up the alert, the eval where one is needed, and an agent that has Lucid file the issue when the alert fires.
  • Versions and a timeline: every rewrite keeps the text before it, and new evidence on a closed issue reopens it as a regression.
  • Hand it to a coding agent: Fix from the issue, or Copy as prompt for any coding agent.
Failures grouped by agent, tool and category, and the issues Lucid files from them.

Trodo Agents

Build agents on top of your own Trodo data. The visual workflow editor wires traces, events and metrics into agents that connect to your apps, your tools and other agents, triggered by a webhook, a schedule, or an alert on an eval. That last one is exactly how self-improving is built.

  • Your data as fuel. Agents read the traces and metrics Trodo already holds.
  • Connected. External tools and MCP servers as dynamic capabilities.
  • Versioned. Draft, publish and roll back workflow versions safely.
Build agents from your Trodo data and connect them to your apps and other agents.

Playground and experiments

Test prompts and models before they ship: run candidates side by side over a dataset, score them with your evaluators rather than vibes, and promote the winner once the numbers back it. Offline testing and production scoring, one definition of "good".

Capabilities

Trodo clusters your runs by what users are actually trying to do and scores how well each cluster is served. You see the jobs your agent handles well, the ones where it falls short, and the ones users keep asking for that it cannot do at all.

Cluster runs by intent, score satisfaction per cluster, and find the gaps in what your agent can do.

Product analytics

The human side of your product: events, funnels, flows, retention and cohorts, on the same user identity as your agent runs. Segment agent quality by product behaviour, or the other way round.

Track product behavior with events, funnels, flows, and retention, all tied to the same user identity as your runs.

Reports

Shared dashboards over your traces: tabs of cards in any chart type, rich text, filters and auto-refresh, with editors and a public link. Every team starts with an Overview of runs, cost, latency, usage and evals, and Lucid builds reports beside the chat.

Build reports from your traces and share them with your team.

Start here

On this page