Skip to main content
A replay run plays every case of a replay set under two arms, prod (the baseline) and proposed (the candidate), once per seed, and a judge scores each playout. These read-only endpoints return the verdicts and the playouts themselves, so you can check a verdict against what the agent actually said and did. The moda replays CLI family calls them.

Base URL and authentication

Like the fix routes, replay routes are tenant-scoped and live under the control-plane base:
They accept the same x-api-key header as the Data API; a key only reads its own tenant. Responses use camelCase field names.

Endpoints

A runId that is not a UUID returns 400; a run that is not this set’s returns 404.

GET /replay-sets/:setId/runs/:runId

Each results[] row:

GET /replay-sets/:setId/runs/:runId/playouts

Returns { runId, playouts[], total, limit, nextCursor, truncated }, ordered by case, arm, then seed. When a run was redispatched, only the newest attempt of each case, arm and seed is listed. Each playout:

GET /replay-sets/:setId/runs/:runId/playouts/:playoutId

Returns one playout (the fields above) with transcript always present. Accepts maxChars.

Transcript

provenance says where a tool output came from: provenance.errorSubtype carries the runtime’s failure subtype when there is one, for example missing_source_state for a call the replay could not serve. Playouts captured before this field existed report replay_runtime with no errorSubtype.