sneak peek

opentracesHub

← back home

The web companion to the CLI. Your team's traces, projects, benches, datasets, and sealed replays become browsable, with the evidence behind every claim attached. This is the real prototype, running in your browser. Drive the whole thing below, or scroll for a page-by-page tour.

opentraces hub
opentraces.ai/hub
OpenTraces Hub — workspace overview

interactive prototype

The live Hub preview is built for desktop. Come back on a larger screen to drive it — for now, here's what each view does.

private alpha

Get your team into the Hub.

Get early access for you and your team, a 15-minute call with the founder to get you set up, plus unlimited trace intelligence credits while the Hub is in preview.

preview

An interactive preview

A live look at what we're building into the Hub, from the trace ledger to the bench that gates your pull requests. Every panel below is the real view, booted on its own, not a screenshot.

the workspaceEvery run your team has made, in one place.
workspace ledger

Every Run In One Ledger

The Hub opens on activity, spend, and the ledger it all stands on: a year of runs as a contribution heatmap, token distribution across projects and models, and the traces underneath. Click a day in the heatmap and the ledger filters to it.

  • heatmap by volume, project or dataset
  • tokens by project, model or harness
  • group the ledger by day, project or dataset
every trace

Find The Session You Half Remember

The full ledger is built for recall. Scrub the day timeline to pin what happened on a given day, narrow by project, dataset or harness, then page through the whole history.

  • day timeline, click a bar to pin it
  • filter by project, dataset or harness
  • counts on every filter value
inside a trace

Read The Whole Run

Replay any run as a conversation: prompts, responses, reasoning, and every tool call inline, with git commits surfaced as checkpoints in the stream. A privacy panel shows what was filtered out before you ever saw it.

  • filter by step type
  • expand any tool call
  • privacy-filtered by default
inside a trace

The Run As A Timeline

The same run, read differently: lanes for user, plan, think, read, exec and write, beside working, local and remote git columns and per-step tokens, duration and context fill. Context-waste and run-health signals sit in the gutter, and the panel says whether the numbers came off the wire or off the record.

  • phase, git and telemetry lanes
  • context-waste flagged inline
  • wire-fidelity or record, stated
inside a trace

One Trace, Sliced Four Ways

A long session is the sum of several trajectories. The Hub cuts it into usable units with a small library of slicers: two deterministic grains cut at capture, two labeled by a small model running locally. Every cut is a gap-free tiling of the same steps.

  • user-turn and change, deterministic
  • milestone and sub-goal, labeled locally
  • the slicer that cut a row travels with it
userplanthinkreadexecwrite
T3 · Hub slicer-feature session (Claude Code)32 steps
trace32 steps · colored by action
S1 user-turndeterministic2 trajectories · new trajectory at each human ask
S2 change-burstdeterministic4 trajectories · edit clusters cut on S1 boundaries
S3 milestonecheap_llm5 trajectories · close on a verified outcome
S4 subgoalcheap_llm2 trajectories · S1 spine + internal agent pivots
projects & pull requestsEach codebase, and the agent work that landed in it.
projects

What This Codebase Has Been Up To

Every project opens on its own pulse: sessions this week, tokens and cost for the month, the session in flight, and a four-week session map with one lane per harness. A flow diagram then follows those sessions to where they ended up, in datasets, in pull request trails, or retained in the bucket.

  • 28-day session map, one lane per harness
  • sessions traced to their destination
  • open PRs, datasets fed, harness mix
pull requests

A Change Gate On Every PR

Open a pull request and the bench's change gate sits above the review: the evidence pack behind it, the guarantees the change touches, and any touched surface that maps to no guarantee, named as surface-drift instead of passing quietly. Below it, the agent trail that built the branch, commit by commit, with intents satisfied, cost, steps and wall time.

  • evidence pack of frozen runs
  • unmapped surfaces named, never silent
  • intent alignment per commit
the benchOne bench per project: does the product still do what you promised?
the bench

Did The Promises Hold?

The atlas is one row per product guarantee, written in product language: the scenario bound to it, the verifier that judges it, the last verdict, and how fresh that verdict is. Every hole is a named state, unbound, stale-run, stale-verifier, no-red-proof, rather than a blank cell.

  • change gate and release gate side by side
  • six named row states, no blanks
  • world signals arrive as a second, later label
the bench

Scenario, Verifier, Disposable Box

Each run pairs a scenario with a specific verifier version, executes it on a leased box, and freezes one record. Skips are listed by name, and a green whose verifier was never shown failing carries red_proof_ref: null instead of being counted as proof.

  • verdict, box and sandbox tier per run
  • skips named, never silently passed over
  • a green with no red proof is marked, not counted
the bench

Rewatch The Run That Proved It

Every frozen run record renders twice: a page for people, with a scrubbable terminal cast, the red proof, the scorecard and the named skips, and the same record as a JSON feed for agents. When no cast exists the record says rewatchable: false rather than faking one.

  • scrub the cast, red proof marked on the track
  • the same record as a feed for agents
  • annotations sit beside the record, never inside it
projectionsEverything you can derive from an app and the work that built it.
projections

Datasets Are Static, Arenas Play

One table for everything projected out of an application and its agentic work. Datasets are static, rows materialized by a workflow; arenas are stateful, where a bench plays a scenario, a capsule plays a trace, and a gym plays a task family. A bench lives in its project, so its row links back home.

  • filter by kind, search by name
  • rows, inbox, pass rate at a glance
  • benches link back to their project
datasets

Review Before You Push

A dataset is local first. Its workflow drops candidate rows into an inbox, you promote the ones worth keeping to staging, and the batch goes to a Hugging Face remote when you say so. Each candidate carries the slice of the source trace it was cut from.

  • inbox, then staging, then push
  • per-row promote or skip
  • slice provenance on every candidate
capsules

Sealed Runs That Travel

A capsule is one failing episode, sealed. Anyone can witness the film straight from the URL with no account, or step through the sealed conversation and trail. The badge is the honest part: today every capsule refuses the reproducible claim and reports floor trust instead of implying a live verdict.

  • witness anonymously, reenacting is gated
  • every dependency mode stated, not assumed
  • swap model, harness or provider to pose a counterfactual
intelligence & opsRead meaning off the whole trace base, and see what it costs.
trace intelligence

Ask Your Trace Base

Ask anything in plain english. Spotlight reads bodies, reasoning and tool calls with an LLM grader, then returns ranked traces with a score, a snippet, and the run's fingerprint.

  • natural language search
  • LLM-graded ranking
  • snippet and fingerprint per hit
metering

Boxes And AI, Metered

Boxes bill by the hour per environment and AI credits are spent per model call. Two meters track the month against its allowance, with every lease and model call itemized underneath and daily lease hours stacked by environment, model spend drawn over them.

  • box credits and AI credits against allowance
  • every lease and model call itemized
  • 7, 30 or 90 days, exportable as CSV
setup & security

What Is Allowed To Leave Your Machines

One page for the whole setup: the Hugging Face identity everything runs under, every machine mirroring its local bucket into a single private remote with its own digest and sync state, and the scrubbing pipeline that has to pass before anything syncs.

  • machines inferred from pushes, no central registry
  • off, basic, recommended or strict presets
  • global redaction strings, applied before review
private alpha

Get your team into the Hub.

Get early access for you and your team, a 15-minute call with the founder to get you set up, plus unlimited trace intelligence credits while the Hub is in preview.