Variables: passing data between steps
Declare what a Python step passes on, read it in a later judge as {{steps.step.key}} or in Python as t.steps.step.key, and use what earlier steps answered, with the rule for what a step can read.
The best evals gather once and judge exactly: a Python step finds the evidence (the tool outputs, the retrieved passages, the failed calls, the expected answer) and every later step reads exactly that instead of the whole trace. Variables are how that evidence travels.
A step has three kinds of value a later step can read:
| What it is | Declared on | Read as | |
|---|---|---|---|
| Outcome | What the step answered: the value its routes decide on | Output & routing | {{steps.check.score}} · t.steps.check.score |
| Context | What a Python step passes on for later steps to use; routes never read it | Settings → Context | {{steps.gather.evidence}} · t.steps.gather.evidence |
| Reasoning | A judge's explanation of its answer | Always there on a judge | {{steps.judge.reasoning}} · t.steps.judge.reasoning |
Declaring context
In a Python step's settings, under Context, click Pass something on and give each value a name and a type: text, number, true / false, list or json. Then return it as the second value:
def evaluate(t):
tools = t.spans.where(kind="tool")
failed = [s for s in tools if s.status == "error"]
evidence = "\n".join(f"[{s.tool_name}] {text(s.output)[:800]}" for s in tools)
return (
"has_failures" if failed else "clean", # the outcome, routed on
{ # the context, passed on
"evidence": evidence, # text
"failed": [s.tool_name for s in failed], # list
"failed_count": len(failed), # number
},
)
Every declared key must be returned, with its type; returning a key you didn't declare is refused too. A name can't be both an outcome and a context key: a step decides with one name and passes on another. Context can be up to 256 KB.
Reading it in a judge
Use {{steps.<step id>.<key>}} anywhere in a prompt. A list renders as numbered lines, so a judge reads it like a note, not a JSON dump:
These tools failed while the assistant worked:
{{steps.failures.failed}}
Final answer:
{{run.output}}
Did the answer tell the user it could not get what it needed (true), or answer as if nothing had failed (false)?The Fields list under the prompt has a group From earlier steps (Tool failures · evidence, Tool failures · failed). Click one to insert it.
Reading it in Python
def evaluate(t):
missing = t.steps.figures.missing # a list another Python step passed on
judged = t.steps.judge.score # an earlier judge's outcome
why = t.steps.judge.reasoning # and its reasoning
return not missing and judged >= 0.8Reading it in a Semantic check
A Grounded check's Context can be From earlier steps: point it at a text, list or json value a Python step gathered, and every claim is checked against exactly that: the retrieved passages, the rows a SQL tool returned.
What a step can read
A step can read another step only if that step runs on every path before it. If Gather is the first step, everything after it can read steps.gather.*. If Judge only runs when a screen was unsure, a step that can be reached without going through the judge can't read steps.judge.*: on some traces it would be empty.
The editor checks this when you save, and says which step and why:
"judge" does not run on every path before "summary": only a step that always runs first can be read.
reads steps.gather.evidnce, but "gather" has no "evidnce". It has: evidence, tools.
A worked example: grounded in the right rows
Gather
A Python step, Rows returned (id rows), reads the SQL tool's spans and passes on the rows as text:
def evaluate(t):
sql = t.spans.where(kind="tool", tool_name="run_sql")
rows = "\n".join(text(s.output)[:2000] for s in sql if s.status == "ok")
return bool(rows), {"rows": rows}Outcome True / false: True → Check the figures, False → Pass (no data used). Context rows (text).
Check
A Semantic check, Figures match the rows: Grounded, Answer = run output, Context = From earlier steps → Rows returned · rows, true at 90%. True → Grounded, False → Judge.
Judge
An LLM judge, Judge the figures, reads {{steps.rows.rows}} and {{run.output}} and scores 0 to 1. Below 0.5 → Invented figures; 0.8 and up → Grounded (judged); between → a person.
The judge never sees the whole trace: only the rows the answer was supposed to come from.
Routing and endings
What a step answers, where each answer goes, number ranges, options, the otherwise route, and the named Pass and Fail endings every walk stops at.
Filters and sampling
Decide which traffic an eval scores: filters on any trace field (one agent, spans by name or kind, runs inside a conversation) and sampling to score a steady share. Plus Filter steps inside an eval.