Semantic checks
True-or-false checks that read meaning without a judge call: Grounded (is every claim supported by the evidence?), Intent (does what you describe hold: answered, repeated, frustrated?) and Matches (is it like a known example?).
A Semantic check answers a yes-or-no question about meaning, using small models Trodo runs for you. There is no provider bill and no prompt to tune: you say what to compare, and where the line is. They are the workhorse of triage, fast enough to run on every trace and sharp enough to settle most of what a judge would otherwise see.
Pick the kind under Check: Grounded, Intent or Matches. All three return one true / false, and all three have a True at line you set.
Grounded
Is every claim in the answer supported by the context? This is the hallucination check.

| Setting | What to pick |
|---|---|
| Answer | The text being checked: the run's output, a span's output, or the conversation's last reply. |
| Context | What the answer should rest on: Its spans (by Kind or Name, their Input, Output or Both, the Newest N), the run input, a field by path (run.metadata.retrieved_context), or something an earlier step passed on (From earlier steps). |
| True at | Claims supported, 50%-100%. |
How it decides: the answer is split into claims, and each claim is checked against every piece of context. A claim with figures is decided by the figures: every number in it must appear in the context (rounding, units like s and ms, percentages as fractions are understood). A claim without figures is decided by a language model trained to tell supported from not supported. The step is true when the share of supported claims reaches the line.
Every result shows the count (8 of 15 claims supported (needs 80%) · 11 spans) and lists the claims that were not supported.
Intent
Does what you describe hold? Write the question in plain words (Was the user's question answered?, Did the user ask the same thing again?, This message expresses frustration.) and pick what it is About.

Trodo reads your wording and picks the right kind of check for it, shown under the box as Reads as:
| Reads as | When your intent is like | How it decides | Needs |
|---|---|---|---|
| A statement about the text | This message expresses frustration. · The reply apologises. | Does the text support the statement? | Any text |
| Whether the question was answered | Was the user's question answered? · Did it answer? | Is each reply relevant to the question before it, and not a refusal or a question back? | A question and an answer: Input and output, or a conversation |
| Whether the question went unanswered | Was the question left unanswered? | The same, turned around | As above |
| Whether the user asked the same thing twice | Did the user repeat themselves? · Did the user ask the same thing again? | Compares each user turn with the earlier ones | A conversation's Whole transcript |
| Whether every request was new | Was every request new? | The same, turned around | As above |
If the text you picked can't answer the question, the editor says what's missing: Needs the question too: pick Input and output.
| About | Level |
|---|---|
| Input and output, Run input, Run output | Run |
| Input and output, Span input, Span output | Span |
| Whole transcript, Last user message, Last reply | Conversation |
True at is a Confidence from 0.05 to 0.99 (0.5 by default). Each result shows how it was read and how confident it was: "Was the user's question answered?" · read as: answered · 0.97 (true at 0.50).
Write statements about the text, not the person: This message expresses frustration reads far more reliably than The user is frustrated. And ask one thing at a time: A or B statements are hard to confirm.
Matches
Is the text close to one of these examples? Paste in examples of a known behaviour: real messages from frustrated users, or the canned reply your agent gives when it gives up. The step is true when the text is close enough to any of them.

| Setting | |
|---|---|
| Text | The run's input or output, a span's input or output (model and agent spans), or a conversation's whole transcript. |
| Examples | As many as you like. Real examples from your own traffic work best. |
| True at | Similarity, 0.5 to 1. |
Matches compares with the vectors Trodo already computes for your traces, so it is nearly instant. A trace scored before its vector is ready waits and is retried.
Which one?
| Question | Check |
|---|---|
| Did it make things up? | Grounded, with the tool outputs or documents as context |
| Did it answer? Did it give up? | Intent, Was the user's question answered? |
| Did the user have to repeat themselves? | Intent on the whole transcript, Did the user ask the same thing again? |
| Is the user frustrated? | Intent on the last user message, This message expresses frustration., or Matches against real examples |
| Is this our known bad reply? | Matches |
| Anything subtler | A judge, behind one of these |
When it errors
| Reason | Meaning |
|---|---|
cross_encoder_missing_input | There was nothing to check, e.g. no span matched the context. Put a Python step before it to route those traces elsewhere (the templates do). |
cross_encoder_config | The check can't read what it was pointed at; the message says what to pick instead. |
cross_encoder_timeout | A very long text took too long. Narrow the context: fewer, newer spans. |
Python steps
Write evaluate(t) over the trace: the fields each level can read, returning an answer and passing context on, the helpers, the libraries, the sandbox and its limits.
LLM judge
A model you choose, on your own provider key, answering your prompt in a declared shape: prompts, template variables, outputs, reasoning, cost, and writing prompts that decide well.