Python steps
Write evaluate(t) over the trace: the fields each level can read, returning an answer and passing context on, the helpers, the libraries, the sandbox and its limits.
A Python step runs your own evaluate(t) on every trace that reaches it. It is exact, instant, and makes no model call, so anything code can decide, a Python step should decide.

def evaluate(t):
# t is the trace: t.run, t.spans, t.conversation, t.meta
# Return the value this step declares in Output:
# True/False for a boolean, a number, or one of the options.
return 'error' not in text(t.run.output).lower()The Fields list under the code shows everything this eval's level can read. Click one to insert it: t.run.output, or t.run.metadata["tier"] (type the key into the box beside Metadata first).
Returning an answer
Return what the step declares on its Output & routing tab:
| Declared | Return | Example |
|---|---|---|
| True / false | True or False | return len(t.spans) < 20 |
| Number, from a to b | A number in that range | return round(t.run.duration_ms / 1000, 2) |
One of none, recovered, unrecovered | One of those strings | return "recovered" |
| Several fields | A dict with every field | return {"score": 0.7, "tone": "ok"} |
To pass something on to later steps, return it second:
def evaluate(t):
tools = [s for s in t.spans if s.kind == "tool"]
evidence = "\n".join(text(s.output)[:800] for s in tools)
return bool(tools), {"evidence": evidence, "tools": len(tools)}Every key you pass on is declared under Context with its type: text, number, true / false, list or json. A later judge reads it as {{steps.<this step>.evidence}}, a later Python step as t.steps.<this step>.evidence. See Variables.
If what you return doesn't match what the step declares (a string where it declared a number, a key missing from the context) the result is errored with the reason. Evals never guess.
The trace object
t is read-only. Every view allows both t.run.metadata.tier and t.run.metadata["tier"]; a missing key is None.
t.run
| Field | |
|---|---|
input, output | What the run was given and what it answered |
agent_name, status, level | The agent (your wrapAgent name), ok / error, … |
error, error_type | The run's error summary |
duration_ms, tokens_in, tokens_out, cost_usd | Rolled up from its spans |
span_count, tool_count, error_count | |
started_at, ended_at | ISO timestamps |
metadata, attributes | What you set with setMetadata |
conversation_id, distinct_id, parent_run_id | |
feedback.rating, feedback.satisfaction, feedback.comment | What your users said, when you send feedback |
t.span and every item of t.spans
| Field | |
|---|---|
kind | llm, tool, retrieval, agent, … |
name, tool_name | |
input, output | |
status, error_type, error_message, status_code | |
model, provider, temperature | Model calls |
tokens_in, tokens_out, cost_usd, duration_ms | |
started_at, ended_at, parent_id, run_id | |
attributes (also metadata) | What you set with setAttribute |
prompt_name, prompt_version_hash | For calls made with a managed prompt |
t.spans is a list you can filter:
t.spans.where(kind="tool") # by kind, name, status, tool_name, model, provider
t.spans.where(name="search", status="error")
t.spans.where(contains="timeout") # in the input or output text
t.spans.where(lambda s: s.duration_ms > 5000)
t.spans.where(kind="llm").count()
t.spans.where(kind="tool").first() # also .last(), .any(), .all(fn), .map(fn), .filter(fn)t.conversation
| Field | |
|---|---|
id, runs_count, first_run_at, last_run_at | Every level can read these |
transcript() | {turns: [{role, content, run_id, at}], truncated, total_turns}: user turns are run inputs, assistant turns run outputs |
runs | Conversation evals only: every run in the thread, oldest first, each with its own spans |
run_verdicts | Conversation evals only: the latest result of every eval on each run in the thread |
What each level can read
t.run | t.spans | t.span | t.conversation | |
|---|---|---|---|---|
| Run eval | ✓ | its spans | ✕ | summary and transcript |
| Span eval | its run | ✕ | ✓ | summary and transcript |
| Conversation eval | ✕ | via t.conversation.runs[i].spans | ✕ | ✓ everything |
Reading a cell marked ✕ raises an error that names it, rather than quietly returning nothing.
t.steps
What earlier steps answered and passed on: t.steps.gather.evidence, t.steps.judge.score, t.steps.judge.reasoning. See Variables.
Helpers and libraries
text(value) is the one way to turn any payload into a string: strings pass through, chat-message arrays are joined, {content} / {text} objects give their text, anything else becomes compact JSON. Always wrap input and output in it: they can be strings, objects or message lists depending on your SDK and framework.
re is available without importing. You can import:
json re string textwrap difflib unicodedata html csv · math statistics decimal fractions random · datetime time calendar zoneinfo · collections itertools functools operator copy dataclasses typing enum heapq bisect · hashlib base64 uuid hmac · urllib.parse · numpy · pandas
The sandbox
Python steps run in an isolated sandbox built to run your code on every trace:
- No network. Opening a connection fails at once: an eval never calls out.
- No subprocesses, no files outside a scratch folder.
- 10 seconds per call by default. A call past that is stopped and the result is errored with
timeout. - What a step returns, and what it passes on, can each be up to 256 KB;
inputandoutputfields larger than 256 KB arrive truncated, marked…[truncated: …].
Examples
import re
NUM = re.compile(r"-?\d[\d,]*\.?\d*")
def evaluate(t):
evidence = " ".join(text(s.output) for s in t.spans.where(kind="tool"))
have = {n.replace(",", "") for n in NUM.findall(evidence)}
figures = [n.replace(",", "") for n in NUM.findall(text(t.run.output))]
missing = [f for f in figures if f not in have]
return not missing, {"missing": missing}Declared: True / false; context missing (list).
LIMITS = {"enterprise": 8, "pro": 15}
def evaluate(t):
tier = t.run.metadata.get("plan") or "pro"
return (t.run.duration_ms or 0) / 1000 <= LIMITS.get(tier, 20)Declared: True / false. Needs run.setMetadata({ plan }) in your agent.
def evaluate(t):
runs = t.conversation.runs
errors = sum(1 for r in runs if r.status == "error")
return len(runs), {"errors": errors}Declared: Number from 0 to 1000 on a conversation eval; context errors (number).
When it errors
The result says which step failed and why:
| Reason | Meaning |
|---|---|
python_no_evaluate | The code defines no evaluate. |
python_exception | evaluate raised: the traceback is on the result. |
python_bad_return / output_type_mismatch | It returned something the step doesn't declare. |
context_mismatch | A declared context key is missing, has the wrong type, or an undeclared one was returned. |
python_import_blocked | It imported a module that isn't on the list above. |
python_network_blocked | It tried to open a connection. |
python_scope_error | It read something this level can't, e.g. t.span from a run eval. |
timeout | It ran past its limit. |
Use Test on a few real traces before turning an eval on: a Python step that fails on real data fails there first.
Choosing a step
The five kinds of step an eval is built from: Python, Semantic check, LLM judge, Filter and Human review. What each is for, what it costs, and how it is set up.
Semantic checks
True-or-false checks that read meaning without a judge call: Grounded (is every claim supported by the evidence?), Intent (does what you describe hold: answered, repeated, frustrated?) and Matches (is it like a known example?).