Python steps

Write evaluate(t) over the trace: the fields each level can read, returning an answer and passing context on, the helpers, the libraries, the sandbox and its limits.

A Python step runs your own evaluate(t) on every trace that reaches it. It is exact, instant, and makes no model call, so anything code can decide, a Python step should decide.

A Python step: its code, the context it passes on (failed, retried, evidence, failed_names), and the list of fields it can read.
def evaluate(t):
    # t is the trace: t.run, t.spans, t.conversation, t.meta
    # Return the value this step declares in Output:
    # True/False for a boolean, a number, or one of the options.
    return 'error' not in text(t.run.output).lower()

The Fields list under the code shows everything this eval's level can read. Click one to insert it: t.run.output, or t.run.metadata["tier"] (type the key into the box beside Metadata first).

Returning an answer

Return what the step declares on its Output & routing tab:

DeclaredReturnExample
True / falseTrue or Falsereturn len(t.spans) < 20
Number, from a to bA number in that rangereturn round(t.run.duration_ms / 1000, 2)
One of none, recovered, unrecoveredOne of those stringsreturn "recovered"
Several fieldsA dict with every fieldreturn {"score": 0.7, "tone": "ok"}

To pass something on to later steps, return it second:

def evaluate(t):
    tools = [s for s in t.spans if s.kind == "tool"]
    evidence = "\n".join(text(s.output)[:800] for s in tools)
    return bool(tools), {"evidence": evidence, "tools": len(tools)}

Every key you pass on is declared under Context with its type: text, number, true / false, list or json. A later judge reads it as {{steps.<this step>.evidence}}, a later Python step as t.steps.<this step>.evidence. See Variables.

If what you return doesn't match what the step declares (a string where it declared a number, a key missing from the context) the result is errored with the reason. Evals never guess.

The trace object

t is read-only. Every view allows both t.run.metadata.tier and t.run.metadata["tier"]; a missing key is None.

t.run

Field
input, outputWhat the run was given and what it answered
agent_name, status, levelThe agent (your wrapAgent name), ok / error, …
error, error_typeThe run's error summary
duration_ms, tokens_in, tokens_out, cost_usdRolled up from its spans
span_count, tool_count, error_count
started_at, ended_atISO timestamps
metadata, attributesWhat you set with setMetadata
conversation_id, distinct_id, parent_run_id
feedback.rating, feedback.satisfaction, feedback.commentWhat your users said, when you send feedback

t.span and every item of t.spans

Field
kindllm, tool, retrieval, agent, …
name, tool_name
input, output
status, error_type, error_message, status_code
model, provider, temperatureModel calls
tokens_in, tokens_out, cost_usd, duration_ms
started_at, ended_at, parent_id, run_id
attributes (also metadata)What you set with setAttribute
prompt_name, prompt_version_hashFor calls made with a managed prompt

t.spans is a list you can filter:

t.spans.where(kind="tool")                  # by kind, name, status, tool_name, model, provider
t.spans.where(name="search", status="error")
t.spans.where(contains="timeout")           # in the input or output text
t.spans.where(lambda s: s.duration_ms > 5000)
t.spans.where(kind="llm").count()
t.spans.where(kind="tool").first()          # also .last(), .any(), .all(fn), .map(fn), .filter(fn)

t.conversation

Field
id, runs_count, first_run_at, last_run_atEvery level can read these
transcript(){turns: [{role, content, run_id, at}], truncated, total_turns}: user turns are run inputs, assistant turns run outputs
runsConversation evals only: every run in the thread, oldest first, each with its own spans
run_verdictsConversation evals only: the latest result of every eval on each run in the thread

What each level can read

t.runt.spanst.spant.conversation
Run eval✓its spans✕summary and transcript
Span evalits run✕✓summary and transcript
Conversation eval✕via t.conversation.runs[i].spans✕✓ everything

Reading a cell marked ✕ raises an error that names it, rather than quietly returning nothing.

t.steps

What earlier steps answered and passed on: t.steps.gather.evidence, t.steps.judge.score, t.steps.judge.reasoning. See Variables.

Helpers and libraries

text(value) is the one way to turn any payload into a string: strings pass through, chat-message arrays are joined, {content} / {text} objects give their text, anything else becomes compact JSON. Always wrap input and output in it: they can be strings, objects or message lists depending on your SDK and framework.

re is available without importing. You can import:

json re string textwrap difflib unicodedata html csv · math statistics decimal fractions random · datetime time calendar zoneinfo · collections itertools functools operator copy dataclasses typing enum heapq bisect · hashlib base64 uuid hmac · urllib.parse · numpy · pandas

The sandbox

Python steps run in an isolated sandbox built to run your code on every trace:

  • No network. Opening a connection fails at once: an eval never calls out.
  • No subprocesses, no files outside a scratch folder.
  • 10 seconds per call by default. A call past that is stopped and the result is errored with timeout.
  • What a step returns, and what it passes on, can each be up to 256 KB; input and output fields larger than 256 KB arrive truncated, marked …[truncated: …].

Examples

import json

def failed(s):
    # Tools often report a failure inside an output whose span says ok.
    if s.status == "error" or s.error_message:
        return True
    try:
        out = json.loads(text(s.output))
    except Exception:
        return False
    return isinstance(out, dict) and bool(out.get("error"))

def evaluate(t):
    tools = t.spans.where(kind="tool")
    bad = [s for s in tools if failed(s)]
    if not bad:
        return "none", {"failed": ""}
    names = ", ".join(sorted({s.tool_name or s.name for s in bad}))
    return "unrecovered", {"failed": names}

Declared: One of none, recovered, unrecovered; context failed (text).

import re

NUM = re.compile(r"-?\d[\d,]*\.?\d*")

def evaluate(t):
    evidence = " ".join(text(s.output) for s in t.spans.where(kind="tool"))
    have = {n.replace(",", "") for n in NUM.findall(evidence)}
    figures = [n.replace(",", "") for n in NUM.findall(text(t.run.output))]
    missing = [f for f in figures if f not in have]
    return not missing, {"missing": missing}

Declared: True / false; context missing (list).

LIMITS = {"enterprise": 8, "pro": 15}

def evaluate(t):
    tier = t.run.metadata.get("plan") or "pro"
    return (t.run.duration_ms or 0) / 1000 <= LIMITS.get(tier, 20)

Declared: True / false. Needs run.setMetadata({ plan }) in your agent.

def evaluate(t):
    runs = t.conversation.runs
    errors = sum(1 for r in runs if r.status == "error")
    return len(runs), {"errors": errors}

Declared: Number from 0 to 1000 on a conversation eval; context errors (number).

When it errors

The result says which step failed and why:

ReasonMeaning
python_no_evaluateThe code defines no evaluate.
python_exceptionevaluate raised: the traceback is on the result.
python_bad_return / output_type_mismatchIt returned something the step doesn't declare.
context_mismatchA declared context key is missing, has the wrong type, or an undeclared one was returned.
python_import_blockedIt imported a module that isn't on the list above.
python_network_blockedIt tried to open a connection.
python_scope_errorIt read something this level can't, e.g. t.span from a run eval.
timeoutIt ran past its limit.

Use Test on a few real traces before turning an eval on: a Python step that fails on real data fails there first.

On this page