Changelog

New updates and improvements to Trodo.

October 2026

October 2, 2026 · trodo-node 2.25.0, trodo-python 2.25.0

Upgrade with npm install trodo-node@^2.25.0 or pip install -U "trodo-python>=2.25.0".

  • Prompt links now reach auto-instrumented provider calls even when compile() ran before the run. If the messages sent to OpenAI, Anthropic or another auto-instrumented provider are exactly the compiled messages, the provider span is linked to the prompt version. Exact match only; edited or truncated messages are left unlinked.

October 1, 2026 · trodo-node 2.24.0, trodo-python 2.24.0

Upgrade with npm install trodo-node@^2.24.0 or pip install -U "trodo-python>=2.24.0".

  • Your agent no longer waits for Trodo. wrapAgent / wrap_agent returns the agent's result the moment it is ready; the finished run goes into a queue and a background sender posts queued runs in batches (every second or 50 runs, gzipped). Before, every agent call waited for its upload and for every retry: a team over its rate ceiling could add up to 90 seconds to one call, and then the trace was dropped.
  • Traces survive outages and rate limits. The sender retries on its own for up to 10 minutes and pauses when the rate ceiling is hit, so a trace is lost only when the queue (10,000 runs) is full, the server rejects it, or 10 minutes of retrying did not get it accepted. Drops are counted and reported, and trodo.deliveryStats() / client.delivery_stats() show the queue.
  • Serverless needs nothing extra. On Lambda, Cloud Functions, Cloud Run, Vercel, Netlify and Azure Functions the SDK detects the platform and sends before returning. Elsewhere the process flushes on exit. Choose the mode yourself with delivery: 'background' | 'immediate'.
  • Python: async with trodo.wrap_agent(...) and @trodo.agent on async def. The decorator on a coroutine function used to close the run immediately and record the coroutine object as its output.
  • A long run's spans streamed before it finished are no longer sent a second time with the finished run.

September 2026

September 30, 2026 · trodo-node 2.23.5, trodo-python 2.23.4, browser SDK 2.3.0

Upgrade with npm install trodo-node@^2.23.5, pip install -U "trodo-python>=2.23.4" or npm install trodo@^2.3.0. The script tag picks up the browser changes on its own.

  • Retries back off the way the server asks. When Trodo answers 429 or 503 with Retry-After, the Node, Python and browser SDKs wait that long before retrying, and every retry delay has random jitter, so many clients that failed together do not all retry in the same instant.
  • The browser SDK retries only what can succeed. Network errors, 5xx and a 429 that says when to retry are retried; other 4xx are not, so a bad request is not sent again and again.
  • Fewer requests from the browser. Session state still updates every few seconds inside the page, but a visible tab sends it every 30 seconds instead of every 5, and a hidden tab sends once when it is hidden and once when it comes back. Session boundaries and durations are counted the same way.

September 30, 2026 · Issues rebuilt, and a Failures tab

Issues open on Failures. Assess → Issues now reads Failures | Issues and lands on Failures: every hard failure in your traces, grouped by agent, tool, error type and category (timeout, rate limited, auth, not found, refused, bad request, server error, connection, code error, other). A failure is counted once, at the deepest span that failed. A row opens every occurrence; an occurrence opens its trace with the failing span selected. See Failures.

Find the failure you mean. A period picker with presets and custom ranges (up to 90 days), search across message, span, tool and error type, and Agent, Tool and Category filters, all applied on the server. The table sorts by last seen, count or users, fifty rows a page with Load more.

Charts above the table. Failures over time as stacked lines, split by category, tool or agent; a breakdown by category, error type or agent; the tools that fail most; and the share of runs that failed. From any row: File as issue, or start an alert already filtered to it.

An issue is one problem, written for whoever fixes it. Lucid files issues from failing traces, with a root cause, the evidence and a category: code, prompt, tool, infrastructure, data or model, the kind of fix it needs. The same cause on another tool is the same issue: Lucid attaches the evidence and rewrites the issue rather than filing a second one, and merges duplicates. Issues from before this release carry over. See How issues are filed.

Versions and a timeline. Every rewrite is a new version that keeps the text before it. The timeline shows each version, what evidence arrived and why, and merges. Evidence shows the step and error it hit, and any eval that failed on the same trace. New evidence on a closed issue reopens it as a regression. See Anatomy of an issue.

Hand it off. Copy as prompt copies the issue for any coding agent: the kind of fix, the root cause, and each failure's step and error. Fix still hands it to Claude Code or Cursor, now with the full evidence. See Fixing an issue.

Track issues in one click. Pick a template and an agent, and Trodo deploys the alert, the eval where one is needed, and a triage agent that has Lucid file issues when the alert fires. Templates cover every hard failure at run and span level (failed steps, tool, model and retrieval failures, each error category, failed runs, error rate), latency, and silent failures checked without a judge. Pause and resume them together. See Track issues.

Soft failures are evals now. Loops, tool call storms, retries running out, tools answering empty, replies in the wrong language and users leaving after a failure are eval templates, with no model call. The pattern detectors, the signals behind them, the Issue thresholds settings and Fix something else are retired: thresholds live on alerts, and old links to the thresholds settings open Alerts.

Agents for the issues flow. Four agent templates: triage alerts into issues, triage then open a GitHub issue, triage then hand the issue to your coding agent, and a weekly issue digest. An alert trigger can watch several alerts and runs when any of them fires. GitHub steps use your team's GitHub App, with no token to paste, and a new step hands an issue to the connected coding agent.

MCP. Issues carry their category and version, list_issues filters by category, get_issue_failures returns a whole page with true totals in one call, and get_run_signals / get_span_signals return the issues a run or span is evidence for and what failed in it. Tool names and arguments are unchanged. See MCP.

September 30, 2026 · Alerts, Lucid and Evals for everyone

New alert opens a pop-up. Start from a template (hard failures, tool failures, failed runs, one kind of failure, error rate, slow runs, slow tool, cost spike), Create with Lucid, or Start from scratch. Alerts on runs and spans can filter by Error category, the Failures tab's categories. See Create an alert.

Lucid and the new Evals for every team. The Lucid chat and the rebuilt Evals are on for everyone; no team is waiting on a rollout.

Steadier ingest. Under heavy load the API now answers 503 with Retry-After instead of hanging until a timeout, a burst from one team can no longer slow ingest for everyone, and the database is faster. Browser events that arrive while the database is busy are answered "retry shortly" instead of being rejected, so the browser SDK sends them again rather than dropping them.

users.upsert / upsert_user work with your site id in the auth header, as the SDK READMEs document. They were refused with "Missing required fields". Fixed on the server; no upgrade needed.

Smaller things. The trace view's span icons are plain icons rather than boxed tiles, and an error shows once instead of three times. Alerts, Agents, Evals, Datasets, Prompts, Experiments, Playground and Traces have real browser tab titles. Idle tabs stop polling.

September 30, 2026 · Billing: Traces, Executions and Lucid credit

Three things on the bill. Traces (every run, span and event stored), Executions and Lucid credit. Allowances are unchanged. See Plans and billing.

Executions cover evals, agents and alerts. One per eval step run on a trace, one per agent step that runs (the trigger, a branch not taken and a retry are free), and one per alert firing (checking an alert is free). Issue investigations are no longer billed.

At the limit. Without usage beyond your plan, evals stop scoring and agents are refused before they start, with the reason on the run; a running agent finishes. Alerts always fire. On Pro, turn on Usage beyond your plan in Settings → Billing and evals and agents keep running, with the extra Executions on your weekly overage invoice. Each project's owner gets an email at 80% and at 100%.

Settings → Usage shows it all. Traces and Executions against the allowance, Executions split into Evals, Agents and Alerts with a daily trend, and Lucid spend. A row past its allowance shows how many over, and billed when overage is on. There is no usage banner above every page any more.

Lucid credit, not Lucid messages. Plans include Lucid credit each month ($2 Free, $6 Trial, $30 Pro), and top-ups never expire. A new account's plan credit is there from the start, and grows at once when you upgrade.

September 26, 2026 · Reports and the Overview

Reports are dashboards over your traces. Assess → Reports holds your team's reports: tabs of cards you drag and resize, a time window and filters for the page, auto-refresh, editors, and a public read-only link. Cards are presets, charts you build over runs, spans, LLM and tool calls, conversations, users, sub-agents and eval results, or rich text. Any of 18 chart types, from stacked areas to Sankey and heatmaps. See Reports.

Lucid builds them. Lucid creates and edits reports beside the chat, and Add to report puts any chart from an answer on one of yours.

Every team has an Overview. Observe → Overview is a ready-made report with Overview, Cost, Latency, Usage and Evals tabs. Change it, or reset it to the default. See The Overview.

"New" never leaves an untitled thing behind. A new agent, dataset or playground is a draft until you save it with a name. The agents' Credentials page is gone: connections live in Settings → Integrations and inside each agent.

September 26, 2026 · trodo-node 2.23.4

Upgrade with npm install trodo-node@^2.23.4. Three fixes, all silent failures before:

  • Vercel AI SDK v7 tool spans carry their output, and a tool that throws shows as an error. Current v7 releases report results in a field earlier SDK versions didn't read, so every tool span landed empty and failed tools looked successful. Parallel tool calls in one step also keep their own inputs now.
  • Vercel AI SDK v7 is captured under ESM and Next.js. The integration couldn't load ai in an ES module, so those apps got no model or tool spans and no warning.
  • trodo.init({ disableInstrumentations }) works. The option was only honoured by registerOTel(), so the documented fix for LangChain double counting in Node was ignored.

September 22, 2026 · Evals rebuilt: triage, self-improving, and fixes for your agent

Evals are rebuilt from the ground up, around one idea: decide the easy cases cheaply, and spend a judge only on the hard ones. An eval is now a small graph of steps. Each trace walks it until something can decide: a Python rule settles what a rule can, a small model settles most of what is left, your LLM judge sees only what those couldn't, and a person sees only what the judge was unsure about. You can score 100% of production for a fraction of what judging every trace costs, and each part of the decision is made by the tool best at it. See How triage works.

Five kinds of step. Python, your own evaluate(t) over the whole trace, in a sandbox with no network, numpy and pandas included. Semantic checks, true-or-false checks on small models Trodo runs, with no provider bill: Grounded (is every claim in the answer supported by the tool outputs or documents? figures are checked exactly), Intent (describe it in plain words, was the question answered?, did the user repeat themselves?, and it reads it as the right kind of check) and Matches (is it like one of these examples?). LLM judge on your own provider key. Filter steps. And Human review, which parks a trace in the grading queue for someone on your team.

Evals score runs, spans or whole conversations. A conversation eval reads every turn and is scored again each time the thread grows. Every result shows the path it took and what each step answered: the judge's reasoning, 8 of 15 claims supported, read as: answered · 0.97.

Pass data between steps. A Python step gathers the evidence once and passes it on; a judge reads exactly that as {{steps.gather.evidence}}, Python as t.steps.gather.evidence. See Variables.

33 templates, one search away. New eval opens a pop-up with templates for the checks most agents need, grounded in tool results, RAG faithfulness, tool call succeeded, recovers from failed tools, prompt injection, personal data, valid JSON, latency and cost budgets, user frustration, conversation resolved, and more, each run on real traces before it shipped. Pick one and the create page fills in; judge steps take one model choice for all of them. Or Start from scratch, or Create with Lucid. See Templates.

Every change is a version; one label says what runs. Saving asks for a note and writes the next version; Set as production is a separate, deliberate step, so drafts, tests and proposals never touch live scoring. Restoring writes a new version. Optionally, re-score the last hours, days or runs of history whenever production changes, with a cost estimate first when it would call your judge. See Testing and versions.

Results you can reason about. Verdicts with their paths; a Version picker beside the time window; and Analysis that takes an eval apart: pass rate per version, where runs end, cost, latency, every path and every step. Versions side by side compares every version over the same window, and on the same traces shows how two versions that scored the same traces agree and which way the rest flipped. See Results and analysis.

Mark results wrong, one or many. 👎 on a result, or tick several and Mark N wrong, with what they should have been and why. Lucid opens beside the page with the case already written up: each result, what the eval said, what you said, and what to investigate.

Self-improving evals. Deploy an alert on the eval and an agent that has Lucid improve it when the alert fires, on failures, a falling pass rate, results people marked wrong, or audit disagreements. Lucid reads the evidence, decides whether the eval or the agent is wrong, tries its change on the same traces, checks what else it would flip and whether it breaks anything a person labelled, and records a proposed change, never live until you put it in production. See Self-improving evals.

Fixes for your agent, not just your evals. When the eval was right and the agent wasn't, the proposal says so and writes what your engineers should change, the instruction, the tool, the retry, in plain words. Proposed changes is one feed of everything waiting for a decision.

Lucid builds and improves evals. Create an eval that catches answers which make up numbers, Lucid reads your agent's traces, asks only what it can't see, designs it as triage, tries it on two real traces and creates it switched off. These three should have passed, it proposes the fix. See Evals with Lucid.

Evals in your trace views and search. A run's, span's or conversation's Scores tab lists every new eval's verdict, its spans' and runs' too, and loads with the rest of the trace. Search traces by results: "Went well"=fail, "Answer quality">0.7, HALLUCINATION.

Alerts on evals. Watch failures, pass rate, decisions, errors, waiting items, scores, results marked wrong or audit disagreements: straight from Create alert on an eval's page.

Also in this release. Judge prompts read a conversation as text: {{conversation.transcript}}, {{conversation.last_user}}, {{conversation.last_assistant}}. The evals list pages 50 at a time with server-side search. Routing ranges in the multi-output editor save correctly.

September 9, 2026 · Alerts

Tell Trodo what to watch, and hear about it once. Alerts watch one number about your runs or spans, latency, cost, error rate, tokens, tool calls, users, satisfaction, and notify a Slack channel or a webhook the moment it crosses a threshold you set. Where Issues are the failures Trodo finds on its own, an alert is the one you asked for: an SLA, a budget, a latency target your team agreed on.

One number, on your terms. Scope to a run or a span, choose a measure and how a window of rows becomes one value, a count, a rate, an average, a sum, or a percentile from p50 to p99, then set the threshold and the window. Windows run from a minute to thirty days and roll continuously, so nothing waits for a clock boundary. See Measures and aggregations.

Filter it down to the thing you actually care about. One agent, one tool, one model, one customer. Rules combine with AND or OR across groups, and reach into the metadata you attach to runs and the attributes you attach to spans, so tenant: acme becomes a per-customer alert with its own threshold. A span alert can filter on its parent run's fields too, which makes "slow tool calls, but only in this agent" a single alert. See Filters.

It fires on the crossing, not for the duration. An alert notifies once, when it goes from healthy to breaching, and then stays quiet while the problem persists: an hour-long incident is one message, not sixty. It returns to OK silently and is ready to tell you again the next time it crosses. This is what keeps a channel worth reading.

And it won't fire on one bad row. A rate needs at least five rows in the window before it means anything: one failing run out of one is a 100% error rate, and an alert at 20% would otherwise trip on every quiet minute. The form says so rather than leaving the floor invisible.

Slack, webhooks, or both. Pick a channel from your connected workspace, or post JSON to any HTTPS endpoint: the payload carries the value, the threshold, the sample size and the rendered text. Write your own message with {{value}}, {{threshold}}, {{window}}, {{count}} and {{link}}, and the link opens Traces already filtered to the window that fired. If a notification cannot be delivered the alert says so, so an archived Slack channel stops looking like a healthy alert. See Notifications.

See whether the threshold was right. Every alert keeps its number over time on a chart across 6 hours to 30 days, with your threshold drawn on it and everything past it shaded red: the shape of the problem, at a glance. Below that: how many times it triggered per day over the last month, and a history table of every fire with its value, sample size and delivery result. The fastest way to set a threshold is to create the alert with an extreme one, leave it a day, and read the chart. See Monitoring alerts.

September 9, 2026 · Lucid, side panels, and chat history

Alerts, from the chat. Ask "alert me when checkout p95 goes above 5 seconds over 15 minutes and send it to #eng-oncall" and Lucid proposes it as a card that is the alert form, scope, measure, aggregation, filters, threshold, window, message, prefilled and editable. The destination is chosen on the card: a channel from your connected Slack (or a Connect Slack button right there when the team has none), or a webhook. Approving creates the alert, switched on if you asked, with a link and an on/off control in the chat. Lucid reads your alerts too when a question calls for it: which one fired, what it saw, and what in the traces explains it. See Alerts.

Cards you could not dismiss, and two you could not create. Dismissing any proposal card in the chat failed quietly, and routine and experiment cards could not be saved at all. Both fixed.

The extended-thinking menu in light mode now uses the same surface, border and shadow as every other brand menu.

Alert pages for teams with a space in the name. Opening an alert on a team called "Trodo App" showed the list instead of the alert. Fixed.

Agent mode is a switch. Lucid stays on its own page, with a toggle at the top that moves into the agent workspace. Agents live in there now rather than in the main sidebar, so the two surfaces stop competing.

Side panels line up. Every drawer that opens from the side, a run, a span, Lucid, Lucid's history, now fills the app's content panel exactly, top to bottom, instead of each one landing somewhere slightly different. Lucid's history opens beside the conversation rather than over it.

The top bar keeps its controls. Opening and closing Lucid or a trace drawer on the Traces page used to take the Conversations / Runs / Spans and List / Map controls away for good, until a reload. They stay put now.

Chats load fast. The chat rail shows your last ten conversations instead of your entire history, with a search alongside the title that opens a dialog to search across all of it.

Comment on any run. Annotating a run that carries no conversation id used to fail outright. Not every team sends one, and now it doesn't matter.

September 8, 2026 · Comments on your traces

Leave a comment on any run, span or conversation. Mention a teammate and they hear about it in their Trodo inbox and, when Slack is connected, by direct message with a link straight to it. Comments sit with the work, at the foot of a run's Activity, so there is nothing extra to open, and you can edit or delete your own. The mention list shows each person's email, so two people with the same name are told apart.

Every mention arrives. A notification used to depend on timing: some comments reached Slack and the inbox, and some quietly did not. They are now sent before the comment is confirmed, so if you see it posted, your teammate has it.

One workspace, one team. A Slack workspace connects to a single Trodo team. Trying to connect it to a second team is refused and tells you where it already lives, rather than silently breaking the first team's notifications.

Lucid reads them. When a question is about what your team said, Lucid looks at the comments and shows them as comments, the remark, who left it, and what it was on, rather than a list of ids.

Search finds more, and faster. Comments are searchable now, and so are the metadata and custom attributes you record on a run, so looking for a tenant, a release or something a colleague flagged actually turns it up. Search from anywhere (Ctrl+K) answers in a couple of seconds where it could take most of a minute.

Links land where they should. A table in an answer opens the page it came from with the same filters already applied, and a comment opens the run or span it was left on.

Your picture. Sign in with Google and your photo is used automatically; upload your own from Profile at any time. It shows on your comments and across your team.

September 7, 2026 · Lucid, clearer answers and how it got there

Better answers. Lucid picks the right shape for your question, a table when you asked to see records, a chart when you asked about a trend, and draws it rather than describing it. Figures carry their units and read the same way everywhere, and a question with several parts gets all of them answered.

The product itself, inside the answer. Records come back as the same tables you'd see on Traces, Evals or Issues: clickable, with a link that opens the page already filtered. And a proposed evaluator or experiment arrives as the real form: choose the provider and model, read the prompt, edit it, then approve.

Answers you can check. Comparisons say which window they used, and any claim that something rose or fell is verified against its own figures before you see it. Each Lucid run's trace now shows how it reached the answer.

Lucid remembers. Tell it how you'd like answers presented and it keeps to that. A name you clarify once resolves on its own from then on.

Faster and steadier. Questions needing several statistics are worked out in one pass, deep research waits for the findings it builds on, and Trodo starts serving sooner after a quiet spell.

August 2026

August 31, 2026 · Lucid: deep research, routines, and credit billing

Lucid is the new analytics assistant behind the chat, replacing the previous analyst with one unified chat surface. A Lucid investigation is a durable run that outlives the request that made it: closing the tab doesn't cancel it, a reload replays what you missed, and two people can watch the same run live.

Quick and Deep. Quick answers get the full tool surface: analytics, runs, spans, evals, semantic search, and the web. Deep research is a multi-wave investigation: intake resolves the names you typed against real entities (and asks a clickable clarifying question when they're ambiguous), a plan gate shows the task list and estimated cost and waits for your approval before anything is spent, and a lead agent dispatches investigators wave by wave, revising its plan as findings come back. An Analyze control in the composer turns on extended thinking with a low/medium/high/max depth slider.

Answers you can trust. Every claim is validated against what the run actually saw before it's written: figures must appear in a tool result the run produced, and every cited run, span, or file must exist and belong to your team. Claims that can't be verified are reported as "explored, not established" rather than silently dropped. Answers render charts, tables, and KPIs from real rows, and cited ids are clickable links into the product.

Routines. Schedule research that keeps happening, and knows when to stay quiet. A scheduled run whose findings haven't meaningfully changed since the last delivered report is recorded but not delivered, so a routine never trains you to ignore it. Routines can be edited, disabled, or deleted.

Findings become fixes you approve. When a finding traces to a managed prompt version, Lucid drafts a concrete prompt edit with a diff. Nothing auto-applies: approval creates an unpublished draft version, and publishing stays a separate human act in prompt management.

Credits and your own models. Lucid usage is metered in credits: a plan allowance that refreshes each billing period, plus top-ups (bought via Stripe) that never expire, with a cost breakdown down to the individual model call. Or flip Use your own models and Lucid runs entirely on your own provider keys, with per-role model settings (lead, worker, embeddings) validated at save time.

August 31, 2026 · Slack for agents, reworked

The Slack integration for agent workflows is rebuilt around the organization's workspace connection: connect Slack once and every agent uses it, with the connection status visible right in the workflow editor.

Slack actions are now resource-oriented: agent authors pick people, channels, messages, and files by name from real pickers instead of pasting Slack IDs. Actions carry proper permission checks (message edit and delete are gated to bot-authored messages), errors from Slack come back readable, and the deprecated Stars actions are removed (Slack replaced Stars with Later).

August 25 to 27, 2026 · Prompt management: editor, usage analytics, and trace links; SDK 2.23

  • A Usage page per prompt version, runs, tokens in/out, cost, latency (avg, p50, p95), and error rate, charted over 24h/7d/30d/90d windows, plus the list of agents running that version, the question that decides whether a prompt change is safe.
  • Prompt ↔ trace links: the trace tree shows which prompt version each span ran (identified by immutable version hash, never by a movable label), span panels link to the prompt, and each run summarizes the prompts used inside it.
  • Editor overhaul: inline naming, Configure/JSON tabs for model config, a Text/JSON/Structured output format picker, provider-shaped call snippets, and a reworked diff view. Model config is now open JSON: any provider-specific parameter saves, so reasoning_effort or thinking budgets no longer bounce off an allowlist.
  • SDK 2.23 (Node and Python): captures Vercel AI SDK v7 (whose new telemetry registry otherwise emits nothing), stops LangChain double-counting model calls when a provider instrumentation is also active, renders compile() byte-identically across Node, Python, and the backend, accepts version hashes in get_prompt, and stops reporting instrumentations as active when they can't actually capture your provider version.

August 19, 2026 · Security hardening

  • OTP brute-force protection, five failed attempts invalidates the code and forces a fresh, rate-limited resend.
  • GDPR endpoints authenticated: the data export, delete, and anonymize routes now require an API key and are scoped to the key's team.
  • Strict Content-Security-Policy is enforced on the app, alongside a full set of security headers.
  • Login fixes: the failed-login limiter counts failures (not successes), and a failed login or signup no longer leaves an infinite loading spinner.

August 18, 2026 · Conversation-level evaluations

Evaluators can now score a whole conversation, not just a single run or span, because "did this person's problem actually get solved?" is answered across the thread, not in one turn. The transcript an evaluator judges is built from extracted turns (the same text the transcript UI shows), and a growing thread is re-scored with the newest score shown, bounded by a debounce, a minimum-runs floor, and a re-score ceiling stated openly in the setup panel.

The conversation drawer gains the same Transcript / Activity / Scores tabs as the run drawer, human grading works on conversations, and the MCP surface supports conversation evaluators end to end. Also fixed: the "Every turn" cadence actually runs every turn instead of waiting on the batch interval, and a strict max_turns is no longer overruled by sibling evaluators.

August 16 to 17, 2026 · Issues: fix an issue from the issue you're looking at

  • A Fix button on the issue itself: connect a coding agent, add instructions, watch progress inline, and get the pushed branch back on the issue. The separate Heal page is retired; free-text healing remains in chat.
  • Fix history as a tab: every attempt is recorded, and a failed attempt is handed to the next one so the agent doesn't retry a fix that already didn't hold.
  • Replayable evidence: agents are served failure evidence they can actually replay, and a repository guard stops a fix landing in the wrong repo.
  • Duplicate-issue merging: Trodo proposes merges of duplicate issues (based on cited code location, root-cause similarity, and run overlap) and writes nothing until a person approves.
  • Configurable regression window: choose how long a closed issue stays "the same issue" when it recurs, per team, instead of a hard-coded seven days.

August 15, 2026 · One Observe surface

Events and Sessions are now one page behind a grain switcher, and Runs, Spans, and a new Conversations tab live behind a single Traces entry: conversations contain runs contain spans, so the tabs read as a zoom control. The sidebar is regrouped into Observe / Assess / Develop.

Conversations are sessionised on a 30-minute inactivity gap (so a thread key reused for days no longer renders as one two-day "conversation"), and a conversation opens as a readable chat transcript with the current run highlighted. Also in this pass: a polished trace tree, rebuilt event and session drawers on the run drawer's design language, drag-to-reorder columns on runs and spans, and duration sanity checks so a run that never closed cleanly can't stretch every chart.

August 15, 2026 · User properties from a backend, and SDK 2.18.0

Agent analytics runs on backends, and backend integrations had no honest way to say anything about a user: only who they were. Runs were attributed and the profile behind them stayed empty. Three changes close that.

users.upsert, identity and traits in one call. Create-or-update a user and their properties idempotently, safe to call on every request:

await trodo.users.upsert('user-42', {
  properties:      { plan: 'pro' },
  setOnce:         { signup_date: '2026-01-04' },
  fixedProperties: { last_location_country: 'India' },
});
trodo.upsert_user('user-42',
                  properties={'plan': 'pro'},
                  set_once={'signup_date': '2026-01-04'},
                  fixed_properties={'last_location_country': 'India'})

Every bag is optional, with none of them it still ensures the user exists, which is what you want at signup. wrapAgent and startRun also take a user bag, so a request handler can attribute a run and enrich the profile in one round trip.

Fixed properties are settable from a backend: the ones a backend can honestly know. first/last_location_country and first/last_location_city now accept server writes. device_type, browser_name, os, referrer and utm_* remain browser observations and are declined by name rather than stored, because a backend supplying them is guessing and a wrong guess corrupts attribution reporting. Where a browser and a backend both write a column, the browser wins: it observed the value, the server supplied one.

Setting properties no longer requires identify() first. people.* and set_group used to reject a user who had no prior identified session with IDENTIFY_REQUIRED, telling the caller to "call Trodo.identify() first": an instruction a backend had no way to follow, since it has no browser session to graduate. Properties now create the user if needed. This is a server-side fix and applies to existing SDK versions too.

Zero-code OTLP path. Exporter-only integrations can carry traits as trodo.user.property.*, trodo.user.set_once.* and trodo.user.fixed.* span attributes.

Also in SDK 2.18.0: per-user context and session caches are LRU-bounded (maxCachedUsers / max_cached_users, default 10,000): previously unbounded, which was a slow leak in long-lived API servers with high user cardinality. Sessions are created lazily, so a property write no longer builds a session it never uses.

See Custom Properties and Identification.

August 11 to 13, 2026 · Rebuilt user pages, dark theme fixes, and faster traces

  • User pages rebuilt for agent analytics: Overview, Agents, and Events tabs that load only what you're looking at, with the runs list paged instead of fetched up front.
  • Much faster traces: the span list ships a fraction of the data it used to, and page loads make far fewer round trips.
  • Theme and polish: MUI and antd components now follow the app theme instead of staying light-locked, drawers dim the page behind them, and loading skeletons are shaped like the tables they stand in for.

August 6 to 12, 2026 · Experiments and Datasets

  • Experiments, run a dataset of scenarios against the thing you're testing: a specific prompt version, a specific model, or your own system (results ingested via the SDK/CLI). Every run pins a baseline, the previous completed run on the same dataset, and reports each scenario as regressed, improved, or unchanged, with baseline output shown beside the new one.
  • Analysis: saved table views with basic and SQL filters, list and summary layouts, and analysis charts across runs.
  • Datasets, scenarios now have a permanent identity, so editing a dataset no longer breaks comparability between runs made before and after the edit. A scenario can be a multi-turn conversation, expected output can be a structured rubric rather than a string, and origin, difficulty, and category are filterable, "how do we do on the hard cases" is a filter, not a new dataset.
  • The demo workspace now ships with seeded prompts, datasets, experiments, and issues so every surface has something real to look at.

August 3 to 14, 2026 · Billing: anniversary periods, simpler plans, overage, and retention

  • Anniversary billing: every account's allowance and billing period roll on its own anniversary (contract start for Enterprise, signup or upgrade day otherwise), not on the 1st of the month.
  • A simpler plan lineup: Free, Pro, and Enterprise; the legacy Growth tier is retired. Trials ask for confirmation before ending early, and converting can't cost you your trial.
  • Automatic overage for Pro: a single toggle on the billing page arms metered overage past your included amounts, with rates read live from Stripe so the number on the switch is the number on the invoice.
  • Enforced limits: rate ceilings are actually enforced (sliding window, with Retry-After on refusals), and data past your plan's retention is expired by a daily sweep.
  • Correct invoices: invoice period labels fixed (no more "Aug 5 - Aug 5" or inverted labels), the usage metric customers see is "issue investigations", and a large batch of reconciliation fixes keeps team, org, and Stripe in agreement behind the scenes.

July 2026

July 17, 2026 · Chat-message span input and SDK 2.11.0

LLM span input is now a standard chat-message array, the same messages you send to the model: span.setInput([{ role, content }, …]). The four standard roles (system, user, assistant, tool) plus a context role for RAG / retrieved documents, in any order, with any number of messages per role.

  • Per-role embeddings: Trodo embeds the input as a whole and each role separately (all user messages as one vector, all system as one, …), powering sharper AI Score detectors: system → rule adherence; user → trajectory and echo; context (else tool + assistant) → grounding, contradiction, and factual retention.
  • Tool-augmented grounding: when no explicit context is sent, tool results and prior assistant turns become the grounding source automatically.
  • Auto-instrumented traffic included: spans captured from OpenAI, Anthropic, LangChain, the Vercel AI SDK and other instrumentors already arrive as message arrays and get per-role treatment with no code change.
  • SDK 2.11.0 (Node and Python), ChatMessage/ChatRole types (Node), native list pass-through (Python), and refreshed examples. The previous { system_instruction, context, query } object is retired and now stored as one opaque input.

See Structure LLM input.

June 2026

June 24, 2026 · Self-hosted developer docs

The developer documentation moved off Mintlify to a self-hosted site built on Fumadocs, with a rebuilt structure and a construction-grid layout. Every section is being rewritten from the product and SDK source for accuracy.

June 19 to 22, 2026 · Capabilities rebuilt as Use Cases

The capability model was rewritten from the ground up. The old density-clustering engine is replaced by an input-derived model: Trodo reads the input of each run, discovers the broad task types users invoke, and classifies runs into them in real time as they are embedded.

  • Real-time assignment: new runs are classified into the agent's existing capabilities the moment they are embedded, so counts stay current.
  • Periodic rediscovery: an hourly background job re-forms the capability set for agents that need it (new agents, lots of new runs, a low-confidence backlog, or a stale taxonomy).
  • Per-use-case analysis: each capability gets an LLM-written summary, friction points, worst examples, and recommendations.
  • New Use Cases view: list and map layouts plus a detail drawer; capabilities are clickable from a run's detail page.

See Capabilities.

June 17, 2026 · Workflow versioning and SDK 2.7.0

  • Agent workflow versioning: publishing a workflow now snapshots an immutable version you can view and restore.
  • SDK 2.7.0 (Node and Python): deterministic stateless server sessions (server:{distinctId}) and a trimmed session payload that drops browser-only fields on the server. Backend identity now resolves correctly with stateless sessions and last_visit semantics were fixed.

June 7 to 10, 2026 · Linear integration, SDK 2.5.0 and 2.6.0

  • Linear integration: OAuth connection plus the full set of Linear actions (create, update, comment, search, sub-issues) for workflows.
  • SDK 2.6.0: JSONB input and output support on spans (1 MB limit) and span vector embeddings, so richer payloads are captured and made searchable.
  • SDK 2.5.0, anonymous distinct_id is minted on every agent surface, and the backend resolves identity on every agent ingest path.

June 1 to 12, 2026 · Agents workflow editor polish

Iteration on the new workflow builder: run a node using the output of previous nodes, datetime global filters wired to the Insights operators, first-time-event trigger matching, and inline filter persistence and dropdown fixes.

May 2026

May 28, 2026 · Agents: visual workflow builder

Trodo Agents launched: an n8n-style visual workflow builder that watches your Trodo data and acts on it. Five trigger types (live event, pattern, schedule, webhook, manual), node categories for data, flow, compute, AI, and integration, and a BullMQ-backed execution worker.

  • Integrations: Slack, GitHub, Linear, Jira, and Salesforce actions, plus a generic HTTP node and remote MCP tools. Slack event ingress is fully wired as a trigger.
  • AI agent node: an LLM that reasons and calls tools in a loop, configured with a model sub-node and tool sub-nodes (HTTP, MCP, or code).
  • Self-hosted Python sandbox for the code node.

See Agents.

May 28, 2026 · Model configurations

Bring-your-own model credentials are unified in Integrations across 12 providers (OpenAI, Anthropic, Gemini, Mistral, Fireworks, Groq, DeepSeek, xAI, Azure OpenAI, Vertex AI, Bedrock, and OpenAI-compatible endpoints), with multiple models per credential and a brand-styled model picker. These power Evaluations and AI workflow nodes.

May 26, 2026 · Boards v2 and report layouts

  • Boards v2: an open-canvas dashboard layout: drag, resize, and arrange any report on a grid, with text, media, and quick-link cards.
  • Report layouts, create_report from Lucid now produces a clean, shareable report board, with a default 30-day window.

May 21 to 25, 2026 · Evals UI overhaul and Lucid hardening

  • Evaluations UI: a performance card with a time-series chart, type-aware summary (numeric, boolean, string), window-bound results, and a searchable schema picker. See Evaluations results.
  • Lucid harness: a code-level harness validates every tool call before dispatch, and the analytics tools (flows, insights, funnels, retention, segments, top users, property distributions) were hardened.

May 17 to 18, 2026 · Permissions and billing

  • Role-based permissions are enforced across the app (owner, editor, viewer).
  • Billing: per-unit metering, a quota gate, smart retries, an audit log, and scheduled jobs for prepaid billing, Stripe reconciliation, and usage-warning emails.

May 11, 2026 · UX Signals Dashboard, SDK 2.4.2, and Chat Tool Expansion

UX Signals Dashboard

A dedicated UX Signals dashboard is now available from the main navigation, consolidating the four UX telemetry surfaces the auto-events SDK has been collecting since March into a single coherent view.

Click heatmap: A canvas overlay renders aggregated click coordinates across any page path. Click density is expressed as a heat gradient from cool blue to red. The heatmap can be scoped by date range, device class, and user segment. Selecting any region opens a filtered event list showing the raw clicks within that area.

Scroll depth chart: A horizontal bar chart shows the percentage of sessions that reached each 10 % depth milestone for a given page. A dashed line marks the fold threshold inferred from the median viewport height reported in page_performance events.

Rage click analysis: A ranked table of elements most frequently rage-clicked, keyed by the element_selector property on rage_click events. Each row shows the element path, unique user count, share of sessions, and a daily frequency sparkline. Clicking a row opens the session list filtered to sessions containing at least one rage click on that element.

Form abandonment funnel, form_start, form_submit, and form_abandon events are assembled into a mini-funnel per form identifier. Abandonment rate and median time-to-abandon are displayed alongside the most common field at which users stopped interacting, derived from the field_name property on field_blur events.

SDK 2.4.2: ESM fix, OTel v2, span ID correction, debug flag

Four issues discovered in production after the 2.4.0 release are resolved in this patch.

  • ESM compatibility, Projects using "type": "module" in package.json received require is not defined because the internal loader for the optional @opentelemetry/sdk-node peer dependency used synchronous require(). SDK 2.4.2 replaces it with a dynamic import() guarded by a try/catch, valid in both ESM and CJS environments.
  • OTel v2 Resource constructor: OpenTelemetry API v2 changed the Resource constructor signature. The registerOtel helper now detects the OTel API version at runtime via the VERSION export and branches between the v1 and v2 constructor forms, eliminating deprecation warnings and TypeErrors on @opentelemetry/api@^2.0.
  • Span ID format: A refactor in SDK 2.3.0 inadvertently changed span IDs from 16-character lowercase hex (the OTel wire format) to hyphenated UUIDs, breaking correlation workflows that joined on span_id. SDK 2.4.2 restores the 16-character hex format. Existing rows are not backfilled.
  • Debug flag, A new debug: true option on init() and registerOtel() prints each exported span (name, trace ID, span ID, attributes, duration) and the OTLP endpoint URL to the console. Defaults to false and is not read from environment variables to avoid accidental enablement.

Chat: issues, heal, and cluster tools

Three new tool groups extend the Chat interface to cover the remaining observability surfaces.

  • Issues tools: List open issues, fetch detail for a specific issue (affected runs, severity history, timeline, example spans), filter by detector type or severity, and update status (acknowledged, resolved, muted).
  • Heal tools: Retrieve the root-cause narrative and suggested fix for any issue in the Heal queue. Trigger a manual heal analysis on a specified issue ID and receive results inline.
  • Cluster tools: Enumerate clusters for a given agent, fetch topic labels and member counts, list runs assigned to a cluster, and retrieve the 2D canvas coordinates. Enables queries such as "which cluster has the most failed runs?" or "list the top 10 runs in the onboarding cluster".

May 7, 2026 · Heal Workflows and Semantic Search

Heal: Automated Root-Cause Analysis and Fix Suggestion

The Heal system is now generally available. It adds a structured remediation workflow to the Issues surface, targeting issues that have been open for more than 24 hours or have spiked in severity since their last check.

Heal queue: A "Heal" tab in the main navigation lists all issues currently queued for analysis. Items enter the queue automatically when an issue's severity score crosses the configured threshold for the first time, or re-crosses it after a previous resolution. Operators can also manually enqueue any issue from its detail page.

Root-cause narrative: The Heal pipeline fetches the 20 most recent runs associated with the issue, extracts the relevant spans and tool call inputs/outputs, and sends a structured prompt to a language model asking it to identify the common failure pattern. The model returns a free-prose narrative (typically 3 to 5 sentences) stored in the issue detail sidebar under "Root Cause".

Suggested fix, The same prompt asks the model to propose a concrete remediation step, a code snippet, configuration change, or process recommendation. The suggestion is displayed in a "Suggested Fix" panel. It is advisory only; no automated changes are made.

Regression guard: After an issue is marked resolved, the Heal system continues monitoring incoming runs. If the issue's per-day event count re-crosses the severity threshold within 14 days, the issue is automatically re-opened and re-enqueued with a "Regression" label. A Slack notification is sent if a webhook is configured in team settings.

Semantic Search

A new Cmd+K (Ctrl+K on Windows/Linux) shortcut opens the global search modal from any page in the dashboard. Semantic search accepts free-form natural language queries and returns results from three index types simultaneously: agent runs, individual spans, and analytics events.

How it works: The search service computes a 1536-dimension embedding of the query text using text-embedding-3-small. A cosine similarity scan over the semantic_search_documents table returns the top candidates across all three document types, ranked by a linear combination of embedding similarity and a recency score that decays over 30 days.

LLM refinement, For queries that produce low average similarity scores (below a confidence threshold defaulting to 0.72), the service sends the query and top-10 candidates to a language model for re-ranking. The model selects the candidates that genuinely answer the query and generates a one-sentence explanation for each. Refined results display the LLM annotation below each result card.

Confidence gating: If after LLM refinement fewer than two results exceed the threshold, a disambiguation panel replaces the result list with up to three suggested reformulations and a free-text field for manual refinement.

Background indexing: New agent runs are indexed within 60 seconds of completion by a BullMQ worker job that concatenates the agent name, input, output, all tool names, and all span names, computes the embedding, and inserts the vector into semantic_search_documents. The title and snippet fields carry a trigram GIN index to support hybrid queries that mix keyword prefix matching with embedding similarity.

May 4, 2026 · Vercel AI SDK Integration and Additional Provider Support

Vercel AI SDK: Zero-Code Instrumentation

Teams using the Vercel AI SDK can now send agent run telemetry to Trodo without importing the Trodo SDK at all. Set two environment variables before the process starts:

TRODO_API_KEY=your_write_key
TRODO_AGENT_NAME=your_agent_name

Trodo's OTLP ingest endpoint accepts the telemetry the Vercel AI SDK emits when experimental_telemetry: { isEnabled: true } is set on generateText, streamText, generateObject, or streamObject calls. The ingest layer maps Vercel AI SDK span attributes to the Trodo span schema: ai.model.id to model name, ai.usage.promptTokens and ai.usage.completionTokens to token counts, ai.toolCall.name and ai.toolCall.args to tool name and input, and ai.response.text to LLM output. Trodo infers a run boundary from the root span of each call tree, grouping multi-step useChat pipelines under a single agent run. A Next.js App Router quickstart is available in the developer docs under Integrations.

Additional provider auto-instrumentation

SDK 2.4.0 ships auto-instrumentation modules for five providers in addition to the Anthropic and OpenAI modules released in March.

  • Amazon Bedrock, InstrumentBedrock / instrument_bedrock wraps @aws-sdk/client-bedrock-runtime and boto3. Intercepts InvokeModel and InvokeModelWithResponseStream calls and records model ID, token counts, latency, and finish reason. Model IDs are mapped to the Trodo pricing table for cost estimates.
  • Google Generative AI, InstrumentGoogle / instrument_google wraps @google/generative-ai and google-generativeai. Token counts are read from usageMetadata on the response. Streaming calls produce a single span covering the full stream duration.
  • Cohere, InstrumentCohere / instrument_cohere wraps cohere-ai and cohere. Both generate and chat endpoints are instrumented. The command ID and finish reason are recorded as span attributes.
  • Mistral, InstrumentMistral / instrument_mistral wraps @mistralai/mistral-node-client and mistralai. Function-call responses record the function name and arguments as span attributes.
  • HTTP fetch, InstrumentFetch / instrument_fetch intercepts fetch (Node 18+) and requests (Python) for providers not covered by a named module. When the target host matches a configurable allowlist of LLM API hostnames, the interceptor parses the JSON request body to extract model name and messages, and the response body to extract token counts.

All modules are opt-in and compatible with the registerOtel OTel bridge from SDK 2.4.0. Spans from auto-instrumented calls carry trodo.auto_instrumented = true in OTLP exports.

May 2, 2026 · OTLP Ingest and SDK 2.4.0 OpenTelemetry Bridge

OTLP Ingest Endpoint

A new OTLP-compatible HTTP ingest endpoint is available at /v1/traces. It accepts trace payloads in both application/x-protobuf (the standard OTLP wire format) and application/json (the OTLP JSON encoding). Any OpenTelemetry SDK or auto-instrumentation library that supports HTTP OTLP export can send traces to Trodo with no Trodo-specific SDK required.

Requests must include Authorization: Bearer <write_key> or the write key in an X-Trodo-Api-Key header. The ingest layer groups all spans sharing a trace ID into a candidate run: the root span (no parent ID) becomes the run record; child spans are stored as agent spans with parent relationships preserved. If a trodo.agent_name resource attribute is present it is used as the agent name, otherwise the OTel service name is used. Span attributes not recognized by the Trodo schema are stored in a metadata JSONB column and accessible in the span detail view under "Raw Attributes". The endpoint is rate-limited at 1,000 requests per minute per write key; payloads exceeding 5 MB are rejected with 413 Content Too Large.

SDK 2.4.0: registerOtel

SDK 2.4.0 ships registerOtel (Node) / register_otel (Python), a new top-level export that bridges Trodo agent run semantics with the OpenTelemetry ecosystem.

registerOtel accepts a mode parameter with three values. "trodo" (the default when no OTel SDK is detected) records spans via the existing wrapAgent / withSpan surface and exports directly to the Trodo ingest API, identical to pre-2.4.0 behavior. "otlp" configures a minimal OTel SDK using @opentelemetry/sdk-node (Node) or opentelemetry-sdk (Python) as a peer dependency and registers the Trodo OTLP endpoint as the exporter. "coexist" attaches Trodo's OTLP exporter to an existing OTel SDK already configured by the host application, so that agent-relevant spans are received by Trodo without displacing the existing exporter.

When called without a mode argument the SDK auto-detects: if OTEL_EXPORTER_OTLP_ENDPOINT is set it defaults to "coexist"; if @opentelemetry/sdk-node is importable as a peer dependency it defaults to "otlp"; otherwise it defaults to "trodo". Regardless of mode, the SDK maps semantic conventions used by OpenAI, Anthropic, LangChain, and Vercel AI SDK telemetry to the Trodo span schema, so spans from third-party instrumentation libraries appear correctly in the run explorer without manual attribute remapping.

Upgrade via npm install trodo-node@2.4.0 or pip install trodo-python==2.4.0.

April 2026

April 27, 2026 · MCP Runless Spans, SDK 2.3.1, and Chat Agent Tools

MCP: Runless Spans Architecture

The MCP server integration has been re-architected to remove the requirement that MCP tool calls belong to an agent run. Previously, a synthetic run was created per Claude conversation session and all tool call spans were attached to it. MCP tool calls now create a single span in agent_spans with no parent run record required, recording the tool name, input, output, duration, and the conversation_id extracted from the MCP request context. A user_id is materialized if the requesting Claude session is linked to a known Trodo user identity.

Per-span billing replaces per-run billing for MCP. Each span carries a cost computed from its duration and a per-millisecond rate. The Token and Cost Analytics page now has a "MCP Tools" tab showing cost by tool name, by day, and by user. A Redis-backed session sweeper that previously detected conversation end to close synthetic runs has been removed, reducing infrastructure dependencies for self-hosted deployments. A migration adds a conversation_id column to agent_spans with a partial index on (conversation_id, tool_name); existing spans have conversation_id = NULL.

SDK 2.3.1: trackMcp / track_mcp

SDK 2.3.1 introduces a dedicated primitive for recording MCP tool calls. trackMcp(options) (Node) and track_mcp(options) (Python) accept toolName, input, output, durationMs, userId, conversationId, and an optional metadata object. The call creates a single span of kind mcp_tool associated with the current active run if one exists, or as a standalone runless span if called outside a wrapAgent context. Spans are added to the same outbound batch queue used by withSpan and flushed every 2 seconds or when the batch reaches 100 items, preventing MCP-heavy workloads from flooding the ingest API with one HTTP request per tool call. Upgrade via npm install trodo-node@2.3.1 or pip install trodo-python==2.3.1.

Chat: eval tools, agent run tools, and cross-domain query chaining

The Chat interface gains access to the full evaluations and agent run query surfaces. Eval tools allow querying evaluator configurations, retrieving results for a specific agent or date range, listing pending human evaluation items, submitting grades, and skipping queue items. Agent run tools cover listing runs with filter and sort, fetching full run detail including all spans and attributes, computing token cost breakdowns, and retrieving tool call analysis reports.

Cross-domain query chaining lets a single question combine results from both domains. For example, "compare the retention rate of users who had failed runs last week vs. those who didn't" triggers list_agent_runs (failed, last week), get_users_for_event to resolve user IDs, then run_retention_query scoped to those users. The final response presents all three results together. For queries where interpreting an intermediate result requires human judgment before the next tool call, Chat presents a confirmation step the user can confirm or redirect.

April 24, 2026 · MCP Server: Tool Catalog, Auth, and Quickstarts

MCP Server

Trodo is now available as an MCP (Model Context Protocol) server, exposing the full analytics and agent observability surface to any MCP-capable LLM client.

Tool catalog: The MCP server exposes over 70 tools across eight categories: Analytics (run_insights_query, run_funnel_query, run_retention_query, run_flow_query, get_property_distribution, list_event_names, list_event_properties, get_users_for_event); Agent Runs (list_agent_runs, get_agent_run, search_agent_runs, get_run_metrics, get_token_cost_breakdown, get_tool_call_analysis, get_agent_feedback_summary); Evaluations (list_evaluators, get_evaluator, get_eval_results, list_pending_human_evals, submit_human_eval_grade, skip_human_eval); Issues (list_issues, get_issue_details, get_issue_members, get_issue_timeline, set_issue_status, get_top_failure_modes, get_top_failing_tools); Clusters (list_use_case_clusters, get_cluster_summary, get_cluster_runs); Users (get_user_profile, get_user_journey, get_user_agent_runs, find_users, get_top_users); Heal (report_heal_branch, get_anomaly_detection); Evaluator Management (backfill_evaluator, toggle_evaluator, test_evaluator).

Authentication: Teams can register an OAuth application from Settings > Integrations. The authorization server at /auth/oauth supports the authorization code flow with PKCE, issuing access tokens (default 1-hour expiry) and refresh tokens (30-day expiry). For simpler setups, a static API key prefixed with trk_ can be generated from Settings > API Keys. Tools are grouped into six permission scopes (analytics:read, analytics:write, agent:read, agent:write, eval:read, eval:write); tokens can be restricted to a subset of scopes. High-fanout tools are rate-limited at 60 calls per minute per token.

Quickstarts: The developer docs include setup guides for Claude Web (project configuration via pasted server URL), Claude Desktop (JSON config file), Claude Code (.mcp.json in the project directory), and Cursor (.cursor/mcp.json). Each quickstart covers authentication, connection verification, and three example prompts demonstrating cross-tool queries.

April 21, 2026 · Additional Issue Detectors and Severity Scoring

Issues: output_quality, ux_rage, and tool_misuse Detectors

Three additional issue detectors ship this week, expanding the automatic monitoring coverage beyond tool failures and conversation breakdowns.

output_quality: Evaluates the final output of each agent run using a language model judge that checks completeness (did the agent address all parts of the request?), factual coherence (are statements internally consistent?), and format adherence (does the output match the expected structure?). Runs scoring below 0.5 are flagged. An issue is created when the flagged rate exceeds 15 % of runs in any 24-hour window, or when the 7-day rolling average drops below 0.4. Runs where the judge's confidence is below 0.6 are placed in the human evaluation queue automatically. The quality threshold, rubric dimensions, and judge model are editable from Settings > Issues.

ux_rage: Monitors agent runs for conversation-level frustration signals: repetition of a semantically equivalent request within the last five turns, explicit negative feedback phrases, very short replies following a long agent response (indicating dismissal), and session abandonment within 30 seconds of the agent's last response. A composite frustration score is computed per run as a weighted sum of signal counts divided by total turns. An issue is created when five or more runs from the same agent exceed the threshold (default 0.45) within a seven-day window. The issue detail view includes a "Frustration Timeline" showing per-turn signal scores across all member runs.

tool_misuse: Identifies cases where a tool is called with structurally valid but semantically incorrect arguments: the tool executes without error, but a lightweight LLM evaluation of the arguments against the tool's description and expected semantics assigns a misuse score above 0.65. When at least three distinct runs contain misuse candidates on the same tool, an issue is created. This detector is particularly useful for identifying systemic prompt engineering problems where the model consistently misunderstands a tool's purpose.

Polymorphic issue members, The agent_issues_members join table now supports three member types: run, span, and event. The output_quality and tool_misuse detectors attach at the span level, giving issue detail views precise attribution rather than implicating the entire run.

Z-score severity scoring

Issue severity is now computed dynamically using a z-score against each detector's 28-day historical event rate: z = (current - mean) / stddev. Issues with z > 2 are rated Elevated; z > 3 are Critical. Issues with fewer than 7 days of history fall back to an absolute threshold. When an issue's severity drops below Elevated and then returns above it within 14 days, a "Regression" badge is applied to the issue and its timeline. The Heal system uses this signal as an automatic re-analysis trigger. The Issues list now sorts by current z-score descending by default.

April 18, 2026 · Issues Framework

Issues: tool_failure, conversation_breakdown, List View, and Status Lifecycle

The Issues system is now live. It continuously monitors completed agent runs and surfaces systematic failure patterns as issues requiring operator attention.

tool_failure detector, Reads the status field on tool-type spans. Spans with status = "error" are tool failure candidates. When three or more runs fail on the same tool within a 24-hour window, an issue is created, including the tool name, agent name, most common error messages (deduplicated by prefix), affected run IDs, first-seen and last-seen timestamps, and a severity score derived from the failure rate relative to total call volume. An initial backfill job runs over the past 30 days of spans so teams see issues populated immediately rather than waiting for new runs.

conversation_breakdown detector: Applies three heuristics to the full transcript of each multi-turn run: loop detection (three or more turns where consecutive agent response similarity exceeds 0.92); contradiction detection (an NLI classifier identifies pairs of agent statements within the same run that are logically inconsistent); and incomplete termination (the run ends with uncertainty markers in the agent's final turn without a preceding tool call that would justify the uncertainty). Runs satisfying at least two of the three heuristics are scored as breakdown candidates. An issue is created when five or more runs from the same agent match within a seven-day window. Offending turn pairs are recorded as span-level members, linking the issue detail directly to the specific spans where breakdown occurred.

Issues list and detail: The Issues section is now in the main navigation. The list shows issue type, agent name, first and last seen dates, severity, affected run count, and status, with filters for type, agent, status, and date range. The detail page includes a severity sparkline, a "Member Runs" table with links to the full run detail page, a 30-day event count bar chart, and a "Detector Output" section showing the raw detector findings. Issues support four statuses, open, acknowledged, resolved, and muted (suppressed for 30 days, re-opening automatically if events continue). Status changes are logged in the timeline with the operator's name and timestamp.

April 13, 2026 · Clustering, Use Cases, and Chat Eval Tools

Clustering: Semantic Run Clustering, UMAP Canvas, and Use Cases View

Agent runs are now automatically clustered by semantic content, grouping runs that address similar user intents and surfacing use-case patterns that would otherwise require manual labeling.

Embedding: After each agent run completes, a background job encodes the run's input as a 1536-dimension embedding using text-embedding-3-small, stored in the agent_runs table.

Centroid assignment: Nightly, a sweep computes the cosine distance from each unassigned run's embedding to all existing cluster centroids for its agent. If the nearest centroid is within the distance threshold (default 0.25), the run is assigned to that cluster. If no centroid is within threshold, a new cluster is seeded with the run's embedding as its initial centroid.

Merge sweep: Nightly, pairs of clusters whose centroids are within 0.15 cosine distance of each other are merged. The merged centroid is the weighted average weighted by run count. Merges are logged for audit purposes.

Topic labeling: After a cluster accumulates 10 or more runs, a language model produces a 2 to 5 word topic label summarizing the common intent of a sample of 10 runs. Labels are refreshed when the cluster's run count doubles or a merge changes its centroid by more than 0.05.

UMAP 2D canvas: UMAP is run nightly with n_neighbors=15 and min_dist=0.05 to produce 2D coordinates stored alongside the cluster assignment. The canvas is a WebGL-accelerated scatter plot where points are sized by run duration and colored by cluster. Cluster centroids are shown as larger filled circles with topic labels. The canvas is pan-and-zoom enabled and supports a merge sweep overlay showing which runs changed cluster assignment, connected to their new centroid with a thin line.

Use Cases view: The Use Cases section under Agent Analytics presents the cluster directory for each agent: topic label, run count, percentage of total runs, and date range. A coverage badge shows what percentage of the agent's total runs have been assigned to a cluster. Unclustered runs are shown separately. Each cluster retains a snapshot history of its centroid and label across nightly sweeps.

Chat: evaluations and agent run tools

Eval tools allow querying evaluator configurations, retrieving results for a specific agent or date range, listing pending human evaluation items, submitting grades, and skipping queue items. Agent run tools cover listing runs with filter and sort, fetching full run detail including all spans and attributes, and computing token cost breakdowns. When a question requires both agent run data and user event data, Chat calls tools from both domains in sequence and presents all results together in a single response.

April 8, 2026 · Evaluations

Evaluations: LLM Judge, Code Evaluator, Human Grading Queue, and Composite

The Evaluations section launches today. It provides a systematic, repeatable way to score the quality of agent outputs across runs.

LLM judge evaluator: Sends each run's input and output to a language model acting as an impartial scorer. A default rubric checks task completion, response quality, and safety. The judge returns a 0 to 1 score and a one-sentence rationale. The judge prompt, model (defaulting to claude-haiku-4-5), scoring rubric, pass threshold (default 0.7), and schedule (on completion, hourly, daily, or manual) are all configurable per evaluator.

Code evaluator: Executes a user-defined JavaScript or Python function against the agent's input, output, and tool call results in a sandboxed environment (isolated V8 context for Node; restricted subprocess for Python) with a 5-second execution timeout. Useful for deterministic checks: JSON schema validation, regular expression matching, assertion of numeric ranges, or custom business-logic rules that do not require a language model.

Human grading queue: When an evaluator is configured with type "human", matching runs are placed into a grading queue. Graders access the queue from the Evaluations section and see one run at a time: the input, the agent's output, any tool calls made, and the grading rubric. They assign a score from 1 to 5 (normalized to 0 to 1) and an optional comment. The queue shows pending item count, estimated grading time from the median of the last 100 graded items, and an item assignment system that prevents duplication.

Composite evaluator: Aggregates scores from two or more child evaluators using configurable weighting. The composite result is Pass if and only if all child evaluators pass.

Results surface: Each evaluator has a results page showing a 30-day time-series chart of average scores, a score distribution histogram, and a table of recent results with the judge's rationale. The run detail page now includes an "Evaluations" panel showing all eval results for that run. Evaluators can be scoped to a subset of runs using event filter expressions and can be backfilled over historical runs up to 90 days.

April 6, 2026 · Boards and Cohorts

Boards: Pinnable Analytics Dashboards

Boards allow teams to pin any combination of Insights charts, Funnel results, Retention curves, and Flow paths to a shared dashboard that persists across sessions. Every chart rendered on these surfaces has a "Pin to Board" button in its context menu. The chart is saved with its full query configuration so that it recomputes live each time the board loads.

Boards use a grid layout with three column widths (1/3, 2/3, full width) and variable row heights; charts can be repositioned and resized by dragging. Boards can be shared with individual team members or made visible to the entire team. Shared boards are read-only for non-owners. A public share link (no login required, read-only) can be generated for external stakeholders. Auto-refresh intervals of 5 minutes, 15 minutes, 1 hour, or manual update all charts in parallel without a full page reload.

Cohorts: Saved User Segments

Cohorts allow teams to define persistent user segments based on property filters and behavioral criteria, then reuse those segments as filters across all analytics surfaces.

Cohorts are defined using a filter builder that supports AND/OR groupings of property conditions and behavioral conditions such as "performed event checkout_completed at least once in the last 30 days" or "did not perform event churned in the last 90 days". Static cohorts capture the user set at definition time. Dynamic cohorts recompute on a schedule (hourly, daily, or weekly) and always reflect the current state of the filter conditions; membership is stored as a materialized list of distinct_id values for fast join performance. Any filter dropdown in Insights, Funnels, Retention, Flows, the Agent Runs list, or Evaluations results now includes a "Cohorts" section. A cohort comparison view shows a Venn diagram of user overlap between any two selected cohorts with counts and percentages.

April 5, 2026 · SDK 2.3.0, LangChain, and LlamaIndex

SDK 2.3.0: Long-Session Run Primitives

SDK 2.3.0 introduces three primitives for managing agent runs that span multiple processes, threads, or HTTP requests.

startRun(options) / start_run(options) creates a run record and returns a run ID, immediately visible in the Agent Runs list with status in_progress. joinRun(runId) / join_run(run_id) returns a run context object for an existing run ID; subsequent withSpan calls within the returned context are attached to the specified run as child spans regardless of whether the calling process started the run, enabling worker processes to contribute spans to a run started by a dispatcher. endRun(runId, options) / end_run(run_id, options) closes an existing run, recording its final status, output, and metadata, and finalizes the token cost summary. wrapAgent is unchanged; the two styles can coexist within the same codebase.

Upgrade via npm install trodo-node@2.3.0 or pip install trodo-python==2.3.0.

LangChain JS, LangChain Python, and LlamaIndex auto-instrumentation

A TrodoCallbackHandler class is available from trodo-node/integrations/langchain (Node) and trodo.integrations.langchain (Python). It implements the respective framework's callback interface and captures LLM calls, tool calls, chain executions, and retrieval operations as Trodo spans without manual withSpan wrapping. Callback events map to Trodo span kinds: handleLLMStart / handleLLMEnd (or their Python equivalents) produce an llm span with model name, prompt, completion, and token counts; handleToolStart / handleToolEnd produce a tool span; handleRetrieverStart / handleRetrieverEnd produce a retrieval span with the query and top-k retrieved document excerpts. For streaming LLM calls, tokens are accumulated and the span is closed with the full completion on end. The handler is compatible with LCEL pipelines in Node and LangChain expression chains in Python.

A TrodoEventListener class is available from trodo.integrations.llama_index (Node and Python). It implements the LlamaIndex event listener interface and maps LLMCompletionStartEvent / LLMCompletionEndEvent, ToolCallEvent, RetrieveEvent, and QueryStartEvent / QueryEndEvent to Trodo span kinds. Register it via Settings.callback_manager.add_event_listener(TrodoEventListener()).

When called within an active wrapAgent context, spans from both integrations are nested under the current run's trace. When called outside any run context, they create a standalone trace.

April 1, 2026 · Chat: Inline Charts and Report Cards

Chat: Inline Chart Rendering and Report Cards

Query results in Chat now render as interactive charts inline in the conversation rather than raw JSON or plain text tables.

The underlying language model selects the chart type appropriate for the result structure: line charts for time series, bar charts for categorical breakdowns, funnel visualizations for multi-step conversion results, and horizontal bar charts for ranked lists. A chart type toggle in the result card allows the user to switch between chart and table views. Aggregate results (total counts, averages, conversion rates) are rendered as compact metric cards with a value, a label, and a delta indicator if a prior period comparison is available.

Each result card has a "Copy data" button that places the underlying result as a tab-separated table on the clipboard. A "Pin to Board" button saves the chart to a Board. Charts are rendered client-side using the same charting library used across the rest of the dashboard, ensuring visual consistency.

March 2026

March 30, 2026 · Chat: Natural Language Analytics Interface

Chat: Natural Language Analytics and Agent Observability Interface

The Chat interface launches today. It provides a conversational entry point to all Trodo analytics and agent observability data, accepting free-form questions and returning results drawn from the full data surface.

Tool-based architecture: Chat operates by injecting callable tools into the underlying language model's context. Each tool corresponds to a Trodo API endpoint spanning analytics queries, agent run exploration, user identity, and session queries. The model selects and calls the appropriate tools to answer each query, then interprets the results and writes a human-readable response. When a question requires combining multiple data sources, the model calls tools in sequence, passing intermediate results forward.

Event catalog injection and fuzzy matching: At the start of each Chat session, the system injects a compressed event catalog listing all event names, their property schemas, and last-seen timestamps, allowing the model to resolve event references by description rather than exact name. When a query mentions an event name that does not exactly match any event in the catalog, fuzzy string matching suggests the closest match. If a unique best match is found, it is used automatically with a note in the response. Property names and their 20 most common observed values are included in the injected catalog so the model can construct valid filter expressions without the user specifying property names exactly.

Session history: Each Chat session is persisted. The left sidebar shows prior sessions with their first message as a title. Sessions can be renamed, pinned, or deleted. Returning to a session restores the full conversation history and all prior tool result context so follow-up questions can build on prior results without re-fetching data.

March 27, 2026 · SDK 2.1.0 GA, Retention, and Flows

SDK 2.1.0 Generally Available: Anthropic and OpenAI Auto-Instrumentation

SDK 2.1.0 exits beta and is generally available for Node.js and Python. All core agent SDK functions, wrapAgent, withSpan, tool, llm, retrieval, feedback: are now stable within the 2.x series.

InstrumentAnthropic (Node) wraps the @anthropic-ai/sdk client at the transport layer. Every messages.create call is automatically wrapped in an llm span recording the model name, the full messages array, the completion, and the usage object including cache read and cache creation token counts. InstrumentOpenAI (Node) wraps the openai npm package, instrumenting chat.completions.create, completions.create, and embeddings.create. createTrodoMiddleware() returns an Express middleware that extracts a user ID from the X-Trodo-User-Id request header (or from a JWT sub claim if jwtSecret is configured) and sets it as the active user identity for any agent runs initiated within the request handler, eliminating the need to manually forward user IDs from the HTTP layer to agent code.

Retention: Cohort Return Rate Analysis

The Retention surface computes cohort return rates: for users who performed a specific starting action in a given time window, what fraction returned to perform a specific return action in each subsequent period? A query requires a starting event, a return event, and a time granularity (day, week, or month). Results are displayed as a retention triangle where rows are cohorts, columns are elapsed periods, and cells show the return rate color-coded from white (0 %) to dark blue (100 %). An average retention curve below the triangle summarizes the overall shape across all cohorts. Breakdown dimensions split the triangle by acquisition channel, plan tier, or any string property.

Flows: Path Analysis and Sankey Visualization

The Flows surface visualizes the sequences of events users take before or after a specified seed event. The query engine computes the N most frequent events that immediately follow (forward) or immediately precede (backward) the seed within the same session, recursively expanding the most frequent next step up to a configurable depth (default 5 steps). Results render as a Sankey diagram where node width is proportional to user count at that step and edge thickness to the users who made that specific transition. Paths that terminate before reaching the configured depth appear as drop-off nodes.

March 23, 2026 · Developer Docs and Groups SDK

Developer Documentation Site

The Trodo developer documentation site launches today, organized into six top-level sections.

Getting Started: A five-minute quickstart covering account creation, write key retrieval, and sending the first event, with a copy-pasteable HTML snippet and one-liners for Node.js and Python server apps.

Event Analytics SDK, Full reference for the JavaScript and Python SDKs: init, track, identify, page, group, reset, alias, auto-events configuration, session management, privacy controls, and the batch ingest API.

Agent Analytics SDK, Full reference for wrapAgent, withSpan, tool, llm, retrieval, feedback, startRun, joinRun, endRun, and all auto-instrumentation modules, with architecture diagrams showing how spans relate to runs and how runs relate to user identity.

Integrations: Setup guides for LangChain (JS and Python), LlamaIndex, Vercel AI SDK, OpenTelemetry OTLP, Amazon Bedrock, Google Generative AI, Cohere, Mistral, and HTTP fetch instrumentation.

MCP Server: Tool catalog reference, authentication guide, and quickstart configurations for Claude Web, Claude Desktop, Claude Code, and Cursor.

Recipes, Ten end-to-end examples: dual-export, cross-service tracing, sub-agent patterns, multi-model routing, evaluation-driven prompt tuning, cohort-based A/B analysis, and more.

Groups SDK: B2B Organization Tracking

group(groupType, groupId, properties) associates the current user with a named organizational entity. groupType names the kind of entity (e.g., "company", "workspace"); groupId is the unique identifier for the specific instance; properties is a flat object of group attributes such as name, plan, industry, and employee count. Calling group() sends a $groupidentify event that creates or updates the group record. All subsequent events from the user are enriched with the group's current properties, accessible in Insights breakdowns and Funnel filters as group.<property_name>. A user can belong to groups of multiple types simultaneously. The Groups section in the main navigation lists all groups of each type with member count, first-seen date, and last-seen date; each group has a detail page showing properties and a list of member users.

March 20, 2026 · Agent Runs Dashboard and Span Waterfall

Agent Analytics: Runs List, Span Waterfall, Run Detail, and Cost Analytics

The Agent Analytics section is now available from the main navigation.

Runs list: A paginated table showing all agent runs with columns for run ID, agent name, status (completed, failed, in_progress, cancelled), start time, duration, input and output token counts, cost estimate, and user identity linked to the user profile if a userId was provided. Column visibility is configurable per-user and persists across sessions. Filtering by agent name, status, date range, user ID, and run duration range is available; active filters are shown as removable chips. A summary card above the table shows total runs, success rate, median duration, median cost per run, and a daily run count sparkline, all updating live as filters change. A "Traces" tab shows all spans across all runs, filterable by span kind, tool name, and duration.

Span waterfall: Each span is rendered as a horizontal bar whose left edge is its start time relative to the run start and whose width is proportional to its duration. Child spans are indented below their parents. Spans are color-coded by kind: LLM calls in blue, tool calls in amber, retrieval in green, generic in grey, error spans outlined in red. Clicking any span bar opens a slide-out panel showing the full span record: name, kind, status, timestamps, duration, parent span ID, model name, token counts, tool name, input/output excerpts, error message, and raw metadata.

Run detail: The detail page shows a metadata card (status, agent name, start/end time, duration, linked user profile), a stats row (token counts including cache read and cache creation, total cost, span count, tool call count, LLM call count), the embedded span waterfall, and raw input/output in syntax-highlighted JSON viewers with copy buttons. A feedback widget allows operators to submit a thumbs up/down rating and optional free-text comment, stored in agent_run_feedback and accessible via the SDK's feedback API.

Token and cost analytics: A dedicated Token and Cost Analytics page shows total spend broken down by agent name, model name, and day. Each agent row shows total input tokens, total output tokens, total cost, cost per run, and cost per successful run. A configurable pricing table covers all major models; per-run cost breakdowns are itemized by model for multi-model runs. Daily and monthly spend alerts send a Slack notification (if a webhook is configured) and show a dashboard banner when crossed.

March 16, 2026 · Auto-Events SDK v2

Auto-Events SDK v2: 15 Captured Event Types, Form Events, Errors, and Page Performance

The auto-events SDK has been rewritten to capture 15 event types automatically with no configuration beyond passing autoEvents: true to init().

Interaction events, Element clicks ($click), rage clicks (rage_click), dead clicks (dead_click), text selection (text_select), and element view (element_view) are captured. Element click events include the CSS selector path, text content truncated to 64 characters, the element's bounding box position, and whether the element is interactive. A rage click is three or more clicks within 600 milliseconds within a 50 px radius; a dead click is a click that produces no DOM mutation within 750 milliseconds. The element_view event fires once per element per page view session via IntersectionObserver. All element-interaction events include an Element Data property group with selector, tag, id, class, text, href, and position.

Form events, form_start fires on first field focus; form_submit fires on form submission; form_abandon fires when the user navigates away from a started but unsubmitted form, including the time spent and the last field with focus. field_focus and field_blur are tracked per field, recording field_name, field_type, and whether the field was empty when focus was lost. form_validation_error fires for each field that fails an HTML5 constraint, recording the field name and constraint type.

JavaScript and network errors: Unhandled exceptions and unhandled promise rejections are captured as js_error events with the error message, stack trace (first 2,000 characters), and throw-site location. Failed HTTP requests (4xx, 5xx, network timeouts, CORS failures) are captured as network_error events by wrapping native fetch and XMLHttpRequest, recording the sanitized URL, HTTP method, status code, and response time. A per-session throttle of 10 errors prevents flooding.

Page performance, A page_performance event is emitted once per page view with Web Vitals metrics: LCP, FID, CLS, TTFB, and FCP, consistent with what Google's web-vitals library reports.

Runtime API, Trodo.autoEvents.enable() / disable() toggle the system at runtime (useful for consent flows). Individual types can be toggled with enableType('rage_click') / disableType('page_performance'). Per-event throttles are overridable at runtime via setThrottle. A beforeCapture(eventType, properties) => properties | null hook provided at init time or set at runtime fires before each event is sent; returning null discards the event, providing a PII scrubbing and conditional suppression hook.

Media events, media_play, media_pause, and media_ended are captured for <video> and <audio> elements, including current playback position, duration, and whether the action was user-initiated or autoplay.

March 14, 2026 · Agent Analytics SDK Beta

Agent Analytics SDK Beta: wrapAgent, withSpan, and Span Types

The Agent Analytics SDK is now available in beta for Node.js (trodo-node) and Python (trodo-python). It provides primitives for instrumenting AI agent workloads and sending structured telemetry to the Agent Analytics section of the dashboard.

wrapAgent(name, fn, options) is the primary entry point. It wraps an async function as a named agent run: the wrapper creates a run record before the function executes and closes it with status and output when the function resolves or rejects. Options: userId, input, metadata. withSpan(name, fn, options) creates a child span within the current active run context, propagated via AsyncLocalStorage (Node) or context var (Python) without requiring explicit context passing. Options: kind (tool, llm, retrieval, or generic), input, output, metadata. Convenience wrappers tool(name, fn), llm(name, fn), and retrieval(name, fn) call withSpan with the appropriate kind. If the wrapped function throws, the span is closed with status: "error" and the error message and stack trace recorded as span attributes; the error is re-thrown so normal error handling is unaffected. feedback(runId, options) records user or operator feedback on a completed run: score (0 to 1), label (positive or negative), comment, userId.

March 9, 2026 · Analytics: Insights, Funnels, Sessions, and Event Catalog

Insights, Funnels, Sessions, Event Catalog, and Privacy Controls

Four analytics surfaces and supporting infrastructure launch this week.

Insights: The core analytics query interface for exploring event volume, unique user counts, and aggregate numeric metrics over time. Metric types: event count, unique users, formula (a mathematical expression combining other metrics), and property aggregation (sum, average, median, percentile, min, or max of a numeric property on matching events). Queries can be scoped to a preset or custom date range with independently selectable time granularity (hour, day, week, month). Any string property can be used as a breakdown dimension, returning one time series per distinct value up to a configurable maximum. Filters use AND/OR groupings of property conditions. Saved queries appear in a sidebar and can be added to Boards.

Funnels: Defines sequences of 2 to 10 event steps. Each step specifies the event name and optional property filters; steps must be completed in order within a configurable conversion window (default 7 days). The funnel displays the number of users who reached each step, the number who converted to the next, the conversion rate, and the median time to convert. Any string property can be used as a breakdown dimension. A "Dropped off users" link at each step boundary opens a filtered user list that can be exported or saved as a cohort.

Sessions page: A paginated table of all sessions with columns for session ID, user identity, first and last event times, duration, page view count, event count, device type, browser, OS, country, and first referrer. The session detail view shows a chronological event timeline; each event expands to show its full property set. Session records include user agent parsing (browser name and version, OS name and version, device type and model) computed server-side at ingest time. UTM parameters and the initial referrer URL are captured at session creation and preserved for the session lifetime for first-touch attribution.

Event catalog: A browsable registry of all event types that have been sent to the project. Events are auto-discovered within 60 seconds of first occurrence. Operators can add human-readable descriptions to any event and any of its properties, visible as tooltips throughout the dashboard and injected into Chat sessions as context. String properties with fewer than 500 distinct values show a frequency table of the top 20 observed values. Events can be marked Hidden (excluded from autocomplete but still queryable) or Verified / Unverified.

Geo enrichment and privacy controls: All incoming events and sessions are enriched with country, city, region, and rounded latitude / longitude derived from the client IP using a locally hosted MaxMind GeoLite2 database. The raw IP is not stored after enrichment. When the DNT: 1 header is present, event ingest is halted for that session. A beforeCapture hook and a hashPII(value) SHA-256 HMAC utility support PII scrubbing at the SDK layer. Per-event-type data retention can be configured from 30 days to 7 years in Settings > Privacy.

March 3, 2026 · JavaScript and Python SDK Beta

JavaScript SDK Beta and Python SDK Beta

The Trodo JavaScript SDK (trodo-node) and Python SDK (trodo-python) are now available in beta.

JavaScript SDK: Available as a self-contained CDN loader snippet (under 500 bytes before gzip) and as an npm package with ES module and CommonJS dual-export supporting tree-shaking. init(writeKey, options) initializes the SDK. track(event, properties) records a named event. identify(userId, traits) links the current session to a known user and merges traits into the user's profile; a UUID anonymousId is generated on first load and persisted in localStorage with a cookie fallback, and historical events are retroactively linked when identify() is called. page(name, properties) records a page view; passing capturePageViews: true enables automatic page view events on every navigation including soft navigations in single-page apps via History API monkey-patching. group(groupType, groupId, properties) associates the current user with an organizational entity. reset() clears the current identity for logout flows. The SDK creates a session on first event and maintains continuity across page loads within a 30-minute inactivity window; session ID, start time, and session number are included on every event. UTM parameters (utm_source, utm_medium, utm_campaign, utm_content, utm_term) are read from the URL query string on initialization and included on every event in the session.

Python SDK, Available on PyPI. track(distinct_id, event, properties), identify(distinct_id, traits), group(distinct_id, group_type, group_id, properties), and page(distinct_id, name, properties) are non-blocking: events are added to an in-memory queue and flushed to the batch ingest API by a background daemon thread every 200 milliseconds or when the queue reaches 100 items, with exponential backoff on API errors. flush() is provided for serverless environments to ensure all queued events are sent before the process exits. TrodoMiddleware for Django and Starlette/FastAPI automatically sets the active user identity from the X-Trodo-User-Id request header and records page events for every incoming request.

Batch ingest API, The /batch endpoint accepts arrays of up to 500 events (max 500 KB) per POST. Each event requires a type (track, identify, page, group, or alias), an ISO 8601 timestamp, a context object identifying the sending SDK, and either anonymousId or userId. Invalid events are rejected individually without failing the entire batch. The endpoint is authenticated via Authorization: Basic <base64(writeKey:)> or the X-Write-Key header. Rate limiting is 10,000 events per minute per write key with 429 Too Many Requests and a Retry-After header on exceedance. Events with a messageId field are deduplicated within a 24-hour window. alias(userId, previousId) creates a permanent link between two user IDs, treating them as the same user in all queries.

On this page