# Build AI agent evals from production conversations

> Export real AI agent conversations from Sentry with the CLI or MCP server and turn them into eval datasets in Braintrust, Langfuse, promptfoo, or Phoenix.

**URL:** https://sentry.io/cookbook/agent-conversations-to-evals/

---

Export your users' real AI agent conversations from Sentry and turn them into eval datasets you score in Braintrust, Langfuse, promptfoo, or Phoenix. Real traffic becomes your ground-truth test set.

**Time:** 30-45 minutes | **Difficulty:** Advanced

## What You'll Learn

- Instrumented an AI agent so every conversation lands in Sentry with model, cost, tokens, and tool calls

- Exported untruncated gen_ai spans through the Sentry CLI and Events API

- Assembled raw spans into eval-ready records with input, output, and telemetry metadata

- Loaded production conversations into Braintrust as an upsert-safe dataset

- Scored real answers with deterministic scorers plus an LLM judge for subjective quality

- Adapted the same pipeline to Langfuse, promptfoo, or Arize Phoenix

**Topics:** AI Observability, Tracing, CLI, Agent Tracing, API

**SDKs:** Next.js, Node.js

## Steps

### 1. Instrument your AI agent with Sentry

If your agent isn't reporting to Sentry yet, add the SDK and enable AI agent monitoring. With the Vercel AI SDK, add `vercelAIIntegration` to your server config and make sure `tracesSampleRate` is above zero. Set `recordInputs` and `recordOutputs` to `true` — the whole point of this recipe is capturing what users actually said and what the model answered, so opt out only on routes that handle sensitive data.

Sentry also has dedicated integrations for OpenAI, Anthropic, LangChain, and other libraries if you're not on the Vercel AI SDK.

### 2. Tag every conversation with a conversation ID

A chat conversation spans many requests, and each request produces its own trace. Calling `Sentry.setConversationId()` in your chat route stamps every AI span from that request with `gen_ai.conversation.id`, so Sentry can group multi-turn conversations together — and so you can export them as complete conversations instead of disconnected fragments.

Use the same ID your frontend already threads through the chat (most chat SDKs, including the Vercel AI SDK's `useChat`, send one).

### 3. Explore your conversations in Sentry

Once traffic flows, open [Explore → Conversations](https://sentry.io/orgredirect/organizations/:orgslug/explore/conversations) in Sentry. Each conversation shows the user's inputs, the model's outputs, which model was used, cost, token counts, and every tool call — and because it's all trace-connected, you can zoom out to the full waterfall from page request to API route to database queries.

This view is your dataset browser: skim it to understand what users actually ask, which conversations thrash, and which models they were routed to. That's the data your evals should run on.

### 4. Export untruncated spans with the Sentry CLI

The Sentry Events API exposes every `gen_ai` span with its full payloads — including complete `gen_ai.request.messages` and `gen_ai.response.text`. The `sentry api` command handles auth for you and works against any `/api/0/` path.

First list conversation IDs (sorted by token volume — the heaviest conversations are the most interesting eval cases), then fetch each conversation's spans. Pages cap at 100 rows, so paginate with `cursor=0::0` — long conversations that exceed one page would otherwise silently lose their tool results and final answer.

The [Sentry MCP server](https://mcp.sentry.dev/) is great for exploring this data interactively from your agent, but its search results are truncated for context efficiency — use the Events API for the full payloads your dataset needs.

### 5. Assemble spans into eval-ready records

Stitch each conversation's spans into one record: `input` is the first user message (parsed from `gen_ai.request.messages`), `output` is the final assistant response, and everything else becomes metadata your scorers can use — model, tokens, cost, tool calls, tool errors, and a link back to the Sentry conversation.

One subtlety that matters: sum tokens and cost from `gen_ai.chat` and `gen_ai.generate_content` spans only. `gen_ai.invoke_agent` spans aggregate their children, so including them doubles every figure. Also filter out your own noise — seeded demo data and smoke tests — by conversation-ID convention before they poison the dataset.

### 6. Load the conversations into Braintrust

Push each assembled record into a Braintrust dataset. Setting `id` to the conversation ID makes inserts upserts — re-running the export syncs new conversations instead of duplicating old ones. Keep the production answer in metadata: your first experiment should score what the agent *actually said* in production, not a re-generation.

You can also log each conversation as a trace in Braintrust's logs so its clustering and topic features work on your production traffic.

### 7. Score real answers with scorers and a judge

Now run an experiment. The task is a passthrough that returns the production output, so you're scoring reality. Write deterministic scorers for anything you can verify with code — if your agent cites entities from your database, groundedness is a lookup, not a judgment call. Telemetry metadata gives you free scorers too: token efficiency and tool success rate come straight from the export.

Save the LLM-as-judge for the one dimension code can't reach — was the answer actually helpful? — and calibrate it periodically against human labels. When you're ready to compare models or prompts, swap the passthrough task for a live call to your agent and run the same dataset through it.

### 8. Swap in the eval stack that fits you

Nothing above is Braintrust-specific except step 6. The pipeline is three separable stages — extract (Sentry Events API), assemble (plain TypeScript), load (your tool's SDK) — so switching platforms means rewriting one function. Solid options:

The fastest way to adapt it: hand the spec to your coding agent. This prompt captures everything this recipe covered, including the traps.

## FAQ

**Do I have to use Braintrust?**

No. Only the load stage is Braintrust-specific — the extract and assemble stages work with any target. Langfuse and Arize Phoenix are open source and self-hostable, and promptfoo runs entirely locally from a JSON file. Step 8 has the mapping for each.

**Should I use the MCP server or the CLI?**

Both work — they hit the same Sentry APIs. The MCP server is ideal for interactive exploration from a coding agent, but its responses are truncated for context efficiency. For bulk export with full message payloads, call the Events API through `sentry api` (or plain HTTP with an auth token).

**What about user privacy and PII?**

Prompts and outputs are only captured when `recordInputs`/`recordOutputs` are enabled, and Sentry supports server-side data scrubbing before storage. Remember the export sends this data to your eval platform too — apply the same scrubbing standards there, or self-host the eval tool.

**Does this replace synthetic eval datasets?**

No — it complements them. Synthetic datasets are guesses about user behavior; production conversations are the ground truth of it. Real traffic tells you what to fix today, synthetic cases protect flows real users haven't exercised yet.

**How many conversations do I need before this is useful?**

Fewer than you'd think. A few dozen real conversations sorted by token volume will surface your agent's actual failure modes — thrashing, tool errors, hallucinated citations — faster than hundreds of synthetic cases.

**Can I compare different models with this?**

Yes, two ways. If your app already routes users to different models, the exported metadata includes the model per conversation — group your experiment by it. To test a model your users haven't touched, swap the passthrough task for a live call to your agent with the candidate model and run the same dataset.

---

*Source: [sentry.io/cookbook/agent-conversations-to-evals/](https://sentry.io/cookbook/agent-conversations-to-evals/)*
