What evals read from your traces

Evals score the traces you already send. The fields that make them sharper (agent names, input and output, conversation ids, span kinds, metadata and feedback) and how to set each from the SDK.

There's nothing to install for evals: they read the traces your agent already sends. But an eval can only check what a trace contains. These are the fields evals lean on, and how to set them.

When traffic is scored

LevelScored when
RunThe run finishes (wrapAgent returns, or endRun).
SpanThe span is recorded.
ConversationA couple of minutes after its latest run finishes, then again each time a new run joins the thread.

Only evals that are on score new traffic, each with its production version.

The fields that matter

FieldUsed bySet it with
Agent nameScoping an eval to one agentThe name you give wrapAgent
Run input and outputJudges ({{run.input}}, {{run.output}}), Semantic checks, Python, transcriptsrun.setInput(…), run.setOutput(…)
Conversation idConversation evals; t.conversation everywhereconversationId / conversation_id on wrapAgent
Span kind and nameSpan evals, span filters, Grounded context, t.spans.where(kind=…)Automatic for framework calls; kind on manual spans
Span input and outputTool and retrieval evidenceAutomatic, or span.setInput(…) / span.setOutput(…)
MetadataFilters (per-customer, per-plan evals), Python, promptsrun.setMetadata({…})
Span attributesSpan filters, Pythonspan.setAttribute(key, value)
Status and errorsTool and model health checksAutomatic; errors you throw are recorded
User feedbackEvals over what users dislikedYour feedback calls; see Feedback
await trodo.wrapAgent('support-bot', async (run) => {       // → agent name
  run.setInput({ query });                                  // → {{run.input}}
  run.setMetadata({ tenant: 'acme', plan: 'enterprise' });  // → filters, t.run.metadata

  const rows = await trodo.withSpan('run_sql', async (span) => {
    span.setInput({ sql });
    const out = await db.query(sql);
    span.setOutput(out);                                    // → evidence for Grounded
    return out;
  }, { kind: 'tool' });                                     // → t.spans.where(kind="tool")

  const answer = await llm.answer(query, rows);             // auto-captured, kind "llm"
  run.setOutput(answer);                                    // → {{run.output}}
  return answer;
}, { distinctId: userId, conversationId: threadId });       // → conversation evals
with trodo.wrap_agent('support-bot', distinct_id=user_id, conversation_id=thread_id) as run:
    run.set_input({'query': query})
    run.set_metadata({'tenant': 'acme', 'plan': 'enterprise'})

    with trodo.span('run_sql', kind='tool') as span:
        span.set_input({'sql': sql})
        rows = db.query(sql)
        span.set_output(rows)

    answer = llm.answer(query, rows)   # auto-captured, kind "llm"
    run.set_output(answer)

Making evals sharper

  • Set the output to the answer the user saw. Not the whole message list, not your internal state: the reply. Judges and Semantic checks read it directly.
  • Put your evidence in spans. Grounded checks an answer against span outputs. A retrieval whose documents never reach a span can't be checked.
  • Use span kinds. tool, retrieval and llm let one eval work across every tool without naming each one: the templates rely on it.
  • Report tool failures honestly. A tool that returns {"error": "…"} inside a successful span looks fine to the dashboards. Throw, set the span's error, or use the Tool call succeeded template, which catches both.
  • Send a conversation id for anything multi-turn. It's what lets an eval ask did the user get what they came for? rather than grading turns in isolation.
  • Put what distinguishes customers in metadata (tenant, plan, region) and you can hold each to its own eval and thresholds.

Limits of what's loaded

A trace's input and output fields are read up to 256 KB each; beyond that they arrive truncated and marked. A run's newest 5,000 spans are loaded, and a conversation's first runs and newest runs up to 200. See Billing and limits.

On this page