How triage works

Decide the easy cases cheaply and send only the hard ones to a judge or a person. When to use which step, and how to design an eval that is both cheap and right.

A judge is the most flexible way to grade an answer, and the most expensive, the slowest, and not always the most exact. Most of your traffic doesn't need it. An empty answer is a fail; a Python rule can see that for free. An answer whose every figure appears in the tool output is grounded; a small model can confirm that in milliseconds. The judge should spend its time, and your money, on what is genuinely hard.

A Triage eval is built that way: cheap steps first, each deciding what it can and routing the rest onward.

What it saves

Move the sliders: your traffic, how much of it the cheap steps settle, and what one call to your judge costs.

Judge every run$1,200600,000 judge calls a month
Triage$18090,000 judge calls a month

Provider spend only: what your model provider bills for judge calls on your key. Steps that are not judges make no model call. Trodo’s own usage is counted separately, one unit per step run.

The saving isn't only money. Traces the screens settle get a verdict in milliseconds rather than seconds, and they are decided exactly: a rule doesn't have off days.

The patterns

Most good evals are one of these shapes.

  1. Screen, then judgePython or a Semantic check decides the clear passes and clear fails. Only the rest reach the judge. The default shape for quality evals.
  2. Gather, then checkPython collects the evidence (tool outputs, retrieved documents, the expected answer) and passes it on. A Semantic check or judge reads exactly that.
  3. Judge, then a personThe judge's clear scores decide. The unsure middle band goes to Human review, and every grade becomes a label the eval learns from.
  4. Classify, then explainA cheap step finds the failures; a judge says why, as one of a few categories. The categories become your report.

The template Grounded in tool results uses three of them at once:

Grounded in tool results: Python gathers the tool results; a Grounded check passes answers whose claims are all supported; a judge scores the rest; its unsure middle band goes to a person.
  1. Gather tool results (Python) collects every tool's output into evidence. A run with no tools passes straight away: there is nothing to be grounded in.
  2. Rests on the tool results (Semantic check, Grounded) splits the answer into claims and checks each against the tool outputs. If 80% or more are supported it passes, with no judge call.
  3. Judge the grounding (LLM judge) reads the same evidence and scores the rest: below 0.5 fails, 0.8 and up passes.
  4. Is it grounded? (Human review) gets the 0.5 to 0.8 band, where the judge was unsure, and a person decides.

Choosing a step

You want to know…UseWhy
Is a field set, a status ok, a number in range, a pattern present?Python or a Filter stepExact and free. Anything code can decide, code should decide.
Is the output valid JSON with the right keys? Under a length? Free of emails and API keys?PythonA parse or a regex is the whole answer.
Does every claim in the answer appear in the tool results or documents?Semantic check → GroundedReads meaning claim by claim, checks figures exactly, no model call.
Was the question answered? Did the user repeat themselves? Is the user frustrated?Semantic check → IntentDescribe it in plain words; it reads it as the right kind of check.
Is this like a failure we already know?Semantic check → MatchesCompare with example texts you paste in.
Is the tone right? Is it in scope? Is it safe? Which of these five failure types is it?LLM judgeNuance and categories. Give it only what the cheaper steps couldn't settle.
We can't tell without a human.Human reviewNever first; always behind a screen, so a person sees a trickle, not a flood.

See Choosing a step for each one in depth.

Simple or Triage

SimpleTriage
ShapeOne step, and a rule for which answers passSteps that route to each other, on a canvas
BuiltIn a form, on the create pageIn the editor, after creating
Good forOne clear question: a judge's score, a Python rule, one Semantic checkAnything where cheap checks can decide part of the traffic

They are the same kind of eval: a Simple eval is just a one-step graph. You can open any Simple eval in Triage and keep adding steps; an eval that is back to one step and a pass/fail reads as Simple again.

How to know the screens are right

A screen that sends a trace straight to an ending saves a judge call, but if the screen is wrong, nobody sees the judge's opinion. Two things keep that honest:

  • The audit slice. A small, fixed share of traffic (0.5% by default) walks the full path even when a screen would have ended it: the screen's decision is recorded, and the walk carries on to the judge. When they disagree, the case goes to the grading queue under Audit disagreements for a person to rule on. Screen accuracy in Analysis shows how often each screen's shortcut disagreed with the full path.
  • People marking results wrong. A 👎 on a result is a label. Enough of them on the same eval, and self-improving proposes a better version.

Anti-patterns

On this page