Who did what
Every eval change, grade, correction and decision is recorded with who made it. Here is where you see it.
Evals decide things about your product, so every change to them is attributable. Two kinds of author appear:
| Author | Shown as | Who |
|---|---|---|
| A person | you / a teammate | Someone on your team, in the product |
| Lucid | authoring | Evals Lucid created, and changes it proposed from a chat, a thumbs-down or self-improving |
What is recorded
| What | Recorded | Where you see it |
|---|---|---|
| Creating an eval, saving a version | Who, when, and the version's note | Versions: author and note on every version |
| Setting a version as production | Every move of the label | Versions; the Pass rate chart marks each one |
| A proposed change | Where it came from, the evidence, what it would flip | Proposed changes |
| Approving, rejecting or reverting a proposal | Who, when, and the reason | The proposal: its status and, for a rejection, the reason. An approval that had to be re-applied is a new version by the person who approved it |
| ๐ / ๐ on a result | Who, the corrected verdict, the reason | Your own marks show filled in on every result; everyone's count toward Marked wrong on Analysis |
| A grade in the queue | Who picked which option, and when | The result's path shows the option; the walk continued from it |
| Confirming a re-score of history | Who confirmed it | Versions: the re-score's status |
| An audit ruling | Who ruled | Screen accuracy on Analysis |
Who can do what
Anyone who can view your team's Trodo can read evals and results. Changing them needs permission to use the app: creating, editing, saving versions, setting production, grading, marking results wrong, deciding proposals and re-scoring history. A view-only teammate can read everything and change nothing.
Why it matters
- Nothing changes silently. An eval's behaviour only changes when a version is set as production, and that is always a recorded act.
- Proposals don't approve themselves. Lucid proposes; a person decides.
- Labels have owners. A correction is one person's judgement. When people disagree with each other you can see it, and so can the next proposal.
What evals read from your traces
Evals score the traces you already send. The fields that make them sharper (agent names, input and output, conversation ids, span kinds, metadata and feedback) and how to set each from the SDK.
Billing, limits and errors
What evals cost in Executions and in your own provider spend, what is never billed, what happens at the plan limit, the limits of what an eval reads, and every state a result can be in.