NOASSERTION · Mixed: devtools frontend is MPL-2.0 (code adapted from Firefox) + BSD (Replay-written files); replay-cli npm package BSD-3-Clause; Chromium fork BSD-3-Clause. The recording/replay backend and the MCP server (https://dispatch.replay.io/nut/mcp) are a proprietary hosted service. Effectively proprietary SaaS with open clients.
language
C++ (instrumented browser runtime), TypeScript (devtools, CLI)
backing
Replay.io, venture-backed startup founded by ex-Mozilla Firefox engineers (Brian Hackett). Funding amount not verified. Product now centred on 'Replay QA' (autonomous app testing), with time-travel recordings as the engine.
stars
722 (as of 2026-10-09)
downloads
4,276/week npm replayio CLI (2026-10-01..07)
latest release
replayio (CLI) 1.9.1, 2026-09-01
first release
2021-06
borrow
Different domain (debugging arbitrary JS runs in the cloud) but the best existing model of what agents want to ask a history.
Borrow from it: take a specific idea, API or format.
the steelman: the best honest case for it
Replay is the strongest proof that agents debug better with time travel than with code reading: it records a whole browser's runtime inputs so the run is exactly reproducible, then hands the agent ~30 MCP tools addressed by execution point — evaluate any expression at any moment, read the stack, see which lines ran, ask why a React fiber re-rendered, read Redux state at a dispatch by path, screenshot the page at a timestamp, and send the human a link to that exact moment. It solved the hard problem (determinism for arbitrary, effectful JavaScript, not just pure Elm-style apps) and the agent-ergonomics problem (summary-first modes, every result carries points the next tool accepts). A fan would say: this is what 'agent-addressable time travel' means in practice, it already works with Claude Code today, and anything scene builds should match its readback depth.
scores
UI as data
0 scene 5
Records arbitrary web apps; the UI stays code.
Agent can drive it
3 scene 5
Rich MCP, but over recordings, not a live instance.
Agent can see it
5 scene 5
Structured state at any point plus screenshots at any timestamp.
Small, targeted edits
1 scene 5
Only retroactive logpoints/evaluation.
Time travel
4 scene 0
Agent can read history and seek to any execution point; no fork or modified replay.
Live data rate
0 scene 5
Post hoc only.
Terminal native
0 scene 5
Pixel graphics
0 scene 4
Teaching
2 scene 3
GetPointLink and shareable recordings help explain bugs; no tour features.
Maturity
3 scene 1
Openness
1 scene 4
Open clients (MPL/BSD), proprietary hosted core and MCP.
Replay.io scene now scene planned
how agents use it
How
Hosted MCP server (dispatch.replay.io/nut/mcp) with ~30 tools, plus agent skills (replay-mcp, replay-cli, replay-playwright). The agent records via `npx replayio record <url>` (a human reproduces, then closes the browser) or a Playwright run, then investigates by recording ID.
Wiring it in
Low: add the MCP server and skills to Claude Code/Cursor/Codex; requires a Replay account (free tier 25 credits/month) and uploading recordings to Replay's cloud.
Seeing the result
Very strong, structured: RecordingOverview, ConsoleMessages, UserInteractions, NetworkRequest, LocalStorage, DescribePoint (variables, executed lines, dependency chain), Evaluate, GetStack, ReactComponentTree at a time, ReduxActions action-state by path, Screenshot at a timestamp, InspectElement.
Small edits
No edits to the recorded program; the only 'writes' are retroactive logpoints and expression evaluation at a point.
History
Seek and read: any execution point or timestamp, for state, stack, DOM, screenshot, React tree, store state. No fork, no replay with modified inputs, no live control.
In short
The most capable agent-facing read-only time travel found: an agent can interrogate any moment of a past run, but cannot change or branch it.
architecture
Runtime replay, not session replay: a modified browser (Chromium fork, also a Node runtime) records the nondeterministic inputs to the JS engine (network, timers, events), so the recording can be re-executed exactly in the cloud. Any 'execution point' can then be reconstructed on demand: evaluate expressions, read stacks and variables, add logpoints retroactively, take screenshots, walk the React fiber tree. Recordings upload to Replay's backend; humans use Replay DevTools in the browser, agents use the hosted Replay MCP server, both against the same recording ID.
Determinism by recording runtime inputs, then re-running: 'When you add a console log to a line that already ran, the messages appear as if they had always been there.'
Execution points are the universal address: every tool that returns events returns points; inspection tools take points.
MCP tools use modes: 'summary' first, then drill-down modes (e.g. ReactRenders summary → waste-rank → component → commit-fibers → fiber-cause).
Framework-aware layers: React tree/renders/exceptions, Redux actions and state at a dispatch, Zustand, TanStack Query, Playwright steps.
Read-only analysis: the recording cannot be changed or branched; 'what if' is via Evaluate at a point.
GetPointLink hands a human a URL to the exact moment the agent is looking at.
performance
No published recording-overhead or replay-throughput numbers found. Analysis is minutes-scale on first open; inherently post hoc.
The first RecordingOverview call on an unanalyzed recording 'can take three to four minutes'; later calls are fast. docs.replay.io ↗
Keep recordings short: a recording that captures just the reproduction is much faster for the agent to analyze. docs.replay.io ↗
Logpoint evaluates an expression at up to 20 hits of a line. docs.replay.io ↗
adoption and upkeep
Adoption
Respected among frontend framework authors; modest CLI usage. Stars are for replayio/devtools. first_release is the npm replayio package creation date, used as a proxy.
Used by: Vercel / Next.js team (testimonial: 'Next.js 13.4 wouldn't have been possible without Replay', Tim Neutkens), Dan Abramov (testimonial)
Maintainability
Actively developed as a company product, but the open repos are peripheral and the company's focus has shifted from devtools to autonomous QA.
Releases: Continuous SaaS deploys; CLI 1.9.1 on 2026-09-01. Product Hunt launches 2026-07-20, 2026-08-17, 2026-09-08 (Replay QA). · Contributors: replayio/devtools 65 (GitHub contributors API) · Recent: replayio/devtools: ~6 commits in last 90 days (open-source frontend mostly quiet); Chromium fork pushed 2026-10-08; MCP docs current. · Bus factor: medium: company-run, closed core; viability tied to one startup's pivot to QA
weaknesses
Proprietary hosted backend; recordings must be uploaded to Replay's cloud (Private Cloud from $1,000+/month, On-Prem from $5,000+/month).
Post hoc only: no live control of a running app, no fork or modified replay.
First analysis of a recording takes 3–4 minutes (docs).
Requires Replay's own browser/runtime; framework tools only for React, Redux, Zustand, TanStack Query, Playwright.
Company focus moved to 'Replay QA' (autonomous testing); the open devtools repo had ~6 commits in 90 days.
Agent-addressable history: read state at a past point, screenshot at a time, shareable link to a moment. scene's planned state_at/screen_at/diff are Replay's DescribePoint/Screenshot for a terminal UI.
What scene would be reinventing
The agent-facing tool shape for history inspection (summary-first, point-addressed, drill-down modes). Replay has already iterated on what an agent needs to read from a past run; scene should copy the ergonomics, not invent them.
The gap it leaves
Replay cannot fork, replay with changes, or drive a live UI; it is web-only, cloud-hosted and closed. scene gives an agent live control plus seek/fork on a local, open, terminal UI whose determinism comes from architecture rather than runtime recording.
What to borrow
One universal address for moments (Replay: execution point; scene: message index / tx id) returned by every history call and accepted by every inspection call.
Tool modes: summary first, then detail, to keep agent context small.
GetPointLink: a human-openable link/command that opens the UI at the exact moment the agent is discussing.
Dependency chain ('why did this change'): scene can answer it exactly from its message log — which message and data-source update produced a node's value.
Retroactive logpoints analogue: evaluate a template/query against the state at every past message.