LLM judge

A model you choose, on your own provider key, answering your prompt in a declared shape: prompts, template variables, outputs, reasoning, cost, and writing prompts that decide well.

An LLM judge sends your prompt, filled in from the trace, to a model you choose, and gets back exactly the fields the step declares. Use it for what needs judgement: whether an answer is relevant, complete, on-brand, in scope, safe; which of five failure types it is.

A judge step's settings: provider and model, parameters, the judge prompt with template variables, and the fields that can be inserted.

The model

Pick a Provider and Model from the providers you added in Settings → Models: OpenAI, Anthropic, Gemini, Azure OpenAI and OpenAI-compatible endpoints. The judge runs on your key: your provider bills you for its calls, and nothing else in an eval does. Params sets temperature (0 by default, judges should be repeatable), max tokens and top-p.

Each judge step picks its own model, so one eval can use a small, fast model for a first opinion and a stronger one only for the cases that reach it.

The prompt

Write what you want decided, and mark where trace data goes with {{…}}:

An assistant answered using tool results.

Question:
{{run.input}}

Tool results:
{{steps.gather.evidence}}

Answer:
{{run.output}}

Score 0 to 1 how well every figure and claim in the answer is supported by the tool results
(1 = all supported, 0 = invented).

You don't write the answer format. Trodo appends an instruction to reply with exactly the step's declared fields as JSON (a number between 0 and 1, one of your options, true or false) and checks what comes back. A malformed reply gets one corrective retry; a second one errors the result.

An optional System prompt sits behind a disclosure under the prompt.

Variables

VariableLevelWhat it holds
{{run.input}}, {{run.output}}Run, span (its run)The run's question and answer
{{span.input}}, {{span.output}}, {{span.name}}SpanThe span's own
{{run.metadata.plan}}, {{span.attributes.region}}Run, spanOne key of your metadata or attributes
{{conversation.transcript}}AnyThe whole thread, one line per turn: user: … / assistant: …
{{conversation.last_user}}, {{conversation.last_assistant}}AnyThe two ends of the thread
{{steps.<step>.<key>}}AnyWhat an earlier step answered or passed on; a list renders as numbered lines
{{steps.<judge>.reasoning}}AnyAn earlier judge's explanation

Click a field in the list under the prompt to insert it. A variable that is empty on a trace renders as nothing; if no variable resolves, the result is errored rather than judged on an empty prompt.

What it returns

Declare the output on Output & routing:

DeclareAsk forRoute
Number 0 to 1"Score 0 to 1 how…"Ranges: fail below 0.5, a person 0.5 to 0.8, pass above
One of"Pick one: complete, partial, missing"One route per option
True / false"Did the assistant comply? Answer true or false."True / False

The judge also returns its reasoning: shown on every result's path, and readable by later steps as {{steps.<judge>.reasoning}}.

Writing prompts that decide well

  • Give it the evidence, not the haystack. Gather what matters in a Python step (the tool outputs, the retrieved passages, the expected answer) and pass exactly that. A judge given a whole trace does worse and costs more.
  • Define the scale. "1 = every claim supported, 0 = invented" beats "rate the quality".
  • Prefer categories to fine scores when you'll act on the category: wrong / unhelpful / tone / other is a report, 0.63 is not.
  • Leave the easy cases out. If a screen has already passed the clean answers, say so: "This answer's figures did not all appear in the tool output."
  • Send the unsure band to a person. Route a middle range of scores to Human review; the grades become labels, and self-improving learns from them.

Cost and speed

A judge step is usually the slowest and the only billed-by-your-provider step in an eval. Its tokens and cost are recorded on every result, summed in Provider spend on Analysis, and estimated before you re-score history. Testing a judge in the editor calls your provider for real.

When it errors

ReasonMeaning
judge_no_model, judge_integration_not_foundNo model picked, or the provider was removed from Settings → Models.
judge_provider_errorThe provider refused the call: a bad key, a retired model, a rate limit. The message is the provider's.
judge_output_invalidTwo replies in a row did not match the declared fields.
judge_template_unresolvedNo variable in the prompt had a value on this trace.
judge_template_scopeThe prompt reads something this level can't.

On this page