What evals read from your traces
Evals score the traces you already send. The fields that make them sharper (agent names, input and output, conversation ids, span kinds, metadata and feedback) and how to set each from the SDK.
There's nothing to install for evals: they read the traces your agent already sends. But an eval can only check what a trace contains. These are the fields evals lean on, and how to set them.
When traffic is scored
| Level | Scored when |
|---|---|
| Run | The run finishes (wrapAgent returns, or endRun). |
| Span | The span is recorded. |
| Conversation | A couple of minutes after its latest run finishes, then again each time a new run joins the thread. |
Only evals that are on score new traffic, each with its production version.
The fields that matter
| Field | Used by | Set it with |
|---|---|---|
| Agent name | Scoping an eval to one agent | The name you give wrapAgent |
| Run input and output | Judges ({{run.input}}, {{run.output}}), Semantic checks, Python, transcripts | run.setInput(…), run.setOutput(…) |
| Conversation id | Conversation evals; t.conversation everywhere | conversationId / conversation_id on wrapAgent |
| Span kind and name | Span evals, span filters, Grounded context, t.spans.where(kind=…) | Automatic for framework calls; kind on manual spans |
| Span input and output | Tool and retrieval evidence | Automatic, or span.setInput(…) / span.setOutput(…) |
| Metadata | Filters (per-customer, per-plan evals), Python, prompts | run.setMetadata({…}) |
| Span attributes | Span filters, Python | span.setAttribute(key, value) |
| Status and errors | Tool and model health checks | Automatic; errors you throw are recorded |
| User feedback | Evals over what users disliked | Your feedback calls; see Feedback |
await trodo.wrapAgent('support-bot', async (run) => { // → agent name
run.setInput({ query }); // → {{run.input}}
run.setMetadata({ tenant: 'acme', plan: 'enterprise' }); // → filters, t.run.metadata
const rows = await trodo.withSpan('run_sql', async (span) => {
span.setInput({ sql });
const out = await db.query(sql);
span.setOutput(out); // → evidence for Grounded
return out;
}, { kind: 'tool' }); // → t.spans.where(kind="tool")
const answer = await llm.answer(query, rows); // auto-captured, kind "llm"
run.setOutput(answer); // → {{run.output}}
return answer;
}, { distinctId: userId, conversationId: threadId }); // → conversation evalswith trodo.wrap_agent('support-bot', distinct_id=user_id, conversation_id=thread_id) as run:
run.set_input({'query': query})
run.set_metadata({'tenant': 'acme', 'plan': 'enterprise'})
with trodo.span('run_sql', kind='tool') as span:
span.set_input({'sql': sql})
rows = db.query(sql)
span.set_output(rows)
answer = llm.answer(query, rows) # auto-captured, kind "llm"
run.set_output(answer)Making evals sharper
- Set the output to the answer the user saw. Not the whole message list, not your internal state: the reply. Judges and Semantic checks read it directly.
- Put your evidence in spans. Grounded checks an answer against span outputs. A retrieval whose documents never reach a span can't be checked.
- Use span kinds.
tool,retrievalandllmlet one eval work across every tool without naming each one: the templates rely on it. - Report tool failures honestly. A tool that returns
{"error": "…"}inside a successful span looks fine to the dashboards. Throw, set the span's error, or use the Tool call succeeded template, which catches both. - Send a conversation id for anything multi-turn. It's what lets an eval ask did the user get what they came for? rather than grading turns in isolation.
- Put what distinguishes customers in metadata (tenant, plan, region) and you can hold each to its own eval and thresholds.
Limits of what's loaded
A trace's input and output fields are read up to 256 KB each; beyond that they arrive truncated and marked. A run's newest 5,000 spans are loaded, and a conversation's first runs and newest runs up to 200. See Billing and limits.