Filters and sampling

Decide which traffic an eval scores: filters on any trace field (one agent, spans by name or kind, runs inside a conversation) and sampling to score a steady share. Plus Filter steps inside an eval.

An eval scores every trace of its level unless you narrow it. Traffic that doesn't match costs nothing and makes no result; Analysis counts it under Not scored · did not match the filter.

Filters

Set them on the create page, or later in Settings → Filters. Changing them makes a new version.

A span eval's filter: span.name is ask.tool.sql.

Click + Add filter, search for a property, pick an operator and a value:

Property typeOperators
TextIs · Is not · Contains · Does not contain · Starts with · Ends with · Matches regex · Is set · Is not set
NumberEquals · Not equal · Greater than · Greater than or equal to · Less than · Less than or equal to · Is set · Is not set
True / falseIs true · Is false · Is set · Is not set

Rules in a group are combined with one AND / OR choice for the group.

What you can filter on depends on the level:

LevelProperties
RunThe run's own fields (agent, status, duration, tokens, cost, error), its metadata keys, and Its spans
SpanThe span's own fields (kind, name, tool, model, status, duration, tokens), its attributes, and its run's fields
ConversationThe thread's fields, and Its runs

One agent

Which agent an eval scores is a filter like any other: Agent name is support-bot. With no agent filter, the eval scores every agent. (The agent name is the name you give wrapAgent.)

Runs that used a certain span

On a run eval, Its spans adds a Spans where … block. Pick Span name or Span kind first to say which span the block is about, then add conditions about that same span:

Spans where Span name is run_sql · Status is error

means runs with a run_sql call that errored. The same way, a conversation eval can filter Runs where …, with span blocks inside.

Per-customer evals

Anything you set with run.setMetadata({ tenant, plan }) becomes a filter: metadata.tenant is acme makes an eval for one customer, with its own versions and thresholds. See What evals read.

Sampling

Sampling decides how much of the matching traffic is scored: Every run, or A share you set as a percentage.

  • The same traces are chosen every time: a trace is always in or always out of one eval's sample, and lowering the share keeps every trace that was already in.
  • Each eval samples independently, so two evals at 10% see different traffic.
  • Traces outside the sample cost nothing and are counted under Not scored · outside the sample.
  • Sampling changes immediately and never makes a version. Re-scoring history ignores it.

Use sampling to cap cost on high-volume agents, and to keep a Human review step's queue at a size people can keep up with.

Filter steps

A Filter step is a rule inside an eval: the same filter builder, answering true when the trace matches, so a walk can branch on a field without writing Python. Status is error → Fail (errored run); otherwise → the rest of the eval.

Use the eval's filters to decide what is scored; use a Filter step to decide which way a scored trace goes.

On this page