> ## Documentation Index
> Fetch the complete documentation index at: https://docs.moda.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Replays

> Tenant-scoped replay result endpoints: run history, one run's per-seed verdicts, and the playout transcripts behind each verdict.

A replay run plays every case of a replay set under two arms, `prod` (the baseline) and `proposed` (the candidate), once per seed, and a judge scores each playout. These read-only endpoints return the verdicts and the playouts themselves, so you can check a verdict against what the agent actually said and did. The [`moda replays` CLI family](/cli/hidden-features#replays) calls them.

## Base URL and authentication

Like the [fix routes](/data-api/fixes), replay routes are tenant-scoped and live under the control-plane base:

```
https://moda.dev/api/tenants/:tenantId/replay-sets/...
```

They accept the same `x-api-key` header as the [Data API](/data-api/overview); a key only reads its own tenant. Responses use camelCase field names.

## Endpoints

| Endpoint | Description |
| - | - |
| `GET /replay-sets` | The tenant's replay sets. |
| `GET /replay-sets/:setId/runs?limit=` | Run history, newest first, at every status (default 20, max 100). |
| `GET /replay-sets/:setId/runs/latest[?runId=]` | The newest run (or `runId`) with per-case aggregates. Kept for polling; use the run endpoint below for detail. |
| `GET /replay-sets/:setId/runs/:runId` | One run at any status: summary, per-case aggregates, every case × arm × seed verdict, and the playout count. |
| `GET /replay-sets/:setId/runs/:runId/playouts` | The run's playouts with their verdicts; optionally with transcripts. |
| `GET /replay-sets/:setId/runs/:runId/playouts/:playoutId` | One playout with its full transcript. |

A `runId` that is not a UUID returns `400`; a run that is not this set's returns `404`.

## GET /replay-sets/:setId/runs/:runId

| Field | Type | Meaning |
| - | - | - |
| `run` | object | `runId`, `status` (`accepted`, `running`, `completed`, `skipped`, `error`), `caseCount`, `prodPassCount` / `proposedPassCount`, the pass rates, `notes` (skip/error reason), the parsed `assistantModel` / `productionModel` / `proposedModel` / `proposedPromptVersion` (null when not stamped), `costUsd` (null when no spend was recorded), `createdAt`. |
| `cases[]` | object\[] | Per case: `prodPass` / `proposedPass` (strict majority over scored seeds; `null` when every seed abstained), mean `prodScore` / `proposedScore`, and the newest seed's rationale per arm. |
| `results[]` | object\[] | One row per case × arm × seed — see below. |
| `playoutCount` | integer | Captured playouts for this run. `0` means there are no transcripts to read. |
| `truncated` | boolean | `true` when the run is larger than the read caps (two verdicts per allowed case × seed pair) and the lists are partial. |

Each `results[]` row:

| Field | Type | Meaning |
| - | - | - |
| `caseId` / `arm` / `seed` | string / string / integer | The playout's coordinates. |
| `pass` | boolean | The judge's verdict. Always `false` when `abstained`. |
| `abstained` / `coverageReason` | boolean / string | The pair hit a replay coverage hole (a tool call the replay could not serve), so the row is not a real verdict and is left out of pass counts. `coverageReason` names the `tool:error_subtype` pairs. |
| `score` | number | Judge score in \[0, 1]. |
| `rationale` | string | Judge rationale. |
| `criteria[]` | object\[] | Per success criterion: `criterion`, `pass`. |
| `promptVersionId` / `promptLabel` | string | The prompt version the arm ran, when pinned. |

## GET /replay-sets/:setId/runs/:runId/playouts

| Query | Default | Description |
| - | - | - |
| `caseId` | — | Only this case. |
| `arm` | — | `prod` or `proposed`. |
| `seed` | — | Only this seed. |
| `include` | — | `transcript` inlines each playout's transcript. |
| `maxChars` | 4000 | Per-field cap for message text and tool JSON, 100–32000. |
| `limit` | 50 | Page size, max 200; max 10 with `include=transcript`. |
| `cursor` | — | The previous page's `nextCursor`. |

Returns `{ runId, playouts[], total, limit, nextCursor, truncated }`, ordered by case, arm, then seed. When a run was redispatched, only the newest attempt of each case, arm and seed is listed. Each playout:

| Field | Type | Meaning |
| - | - | - |
| `playoutId` | string | 64-character hex ID for the playout endpoint. |
| `caseId` / `arm` / `seed` | string / string / integer | `seed` is `null` for a playout captured before seeds were recorded whose judge rationale does not identify exactly one result row. |
| `status` / `nodeCount` | string / integer | Capture status and size. |
| `sourceConversationId` / `replayMode` | string | The production trace the case came from, and how its source state was built. |
| `judge` | object | The judgement stored with the playout: `pass`, `score`, `rationale`. |
| `verdict` | object \| null | The stored result row for this case, arm and seed (same fields as `results[]`); `null` when `seed` is `null`. |
| `transcript` | object | Only with `include=transcript` — see below. |

## GET /replay-sets/:setId/runs/:runId/playouts/:playoutId

Returns one playout (the fields above) with `transcript` always present. Accepts `maxChars`.

### Transcript

| Field | Type | Meaning |
| - | - | - |
| `scenario` | string | The case scenario the simulated user played. |
| `events[]` | object\[] | In order. `type: "message"` has `role` (`user` is the simulated user), `turn`, `content`. `type: "tool_call"` has `toolName`, `status` (`success` / `failure`), `retryIndex`, `input`, `output`, `errorMessage`, `errorClass`, and `provenance`. |
| `finalAnswer` | string | The agent's last reply. |
| `judgeCriteria[]` | object\[] | `criterion`, `pass` as the judge recorded them on the playout. |
| `truncated` | object | `nodes` is `true` when the playout was longer than 1000 steps; `fields` counts values clipped at `maxChars`. A clipped JSON value becomes `{ truncated, originalChars, preview }`. |

`provenance` says where a tool output came from:

| `source` | Meaning |
| - | - |
| `replay_runtime` | The replay runtime's compiled mock over the case's recorded source state. |
| `agent_tool_simulator` | An LLM simulator grounded in the tenant's recorded tool receipts. `simulator` has `engine`, `model`, `recordedNeighbors`. |
| `counterfactual_reconstruction` | A counterfactual reconstruction. `counterfactual` has `reconstructionId`, `evidenceTier`, `failurePolicy`. |
| `generic_handler` | The generic fallback handler. |

`provenance.errorSubtype` carries the runtime's failure subtype when there is one, for example `missing_source_state` for a call the replay could not serve. Playouts captured before this field existed report `replay_runtime` with no `errorSubtype`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.