Out of the loop.

Open your AI agents' black box. The waste is between your agents, not inside them.

A retrieval that fires twice on the same query. A planner that hands off and gets the same thing back. Boxdawn reads the trace and finds the work your agents already did. Detection runs today, in your browser or on your machine.

$ pip install boxdawn

runs locally · no account to try · open source · LLM-judge opt-in

The Boxdawn report for a public Claude Code trace, showing a waste ratio of 65.95 percent, $1.67 of waste against $2.52 analyzed, and the amount attributed to each detector.
The hosted analyzer on a trace anyone can download: davanstrien/agent-race-traces, claude-code.jsonl, CC-BY-4.0. Dollar figures come from a per-model price table with a published source for every rate; the byte ratio never touches a price. Run the same file and you get the same report.

17,881 traces · four public corpora · MIT licensed

28 Claude Code sessions · 6,780 Toolathlon trajectories across 22 frontier models · 10,056 Exgentic agent-LLM traces · 1,017 Claude Code sessions we did not collect


91.9%
Context resend
2026-08 · davanstrien/agent-race-traces (CC-BY-4.0) · input tokens re-sent (byte-exact)
31.7%
LLM-judge precision
2026-08 · unified n=48 · CI [23.1%, 41.0%] · Amendment v2
3/4
Framework support
2026-08 · Anthropic · LlamaIndex · OpenAI Agents SDK PASS on real workload

the problem

None of these tools is watching the gap.

APMs watch services. Trace tools watch a single agent. Gateways watch calls. None of them watches what an agent does after a step goes wrong, which is to try again: re-run the call, re-read the file, re-send the context. The retry is what reaches the bill, and it happens in the gaps between them.

Boxdawn labels the waste it finds, and one of those labels is an agent that received the same error twice, usually because it re-ran the call without reading the message. A retrieval that fires a second time on the same query. A planner that hands off, gets the same output back, and hands off again. None of these show up inside a single span, and all of them are paid for.

input tokens sent to the model2,238,628
already sent before · 2,056,739 · across 1,720 sends
new · 181,889
The hosted analyzer on davanstrien/agent-race-traces, claude-code.jsonl, CC-BY-4.0, run 2026-08-31. The axis is input tokens: resent input tokens over total input tokens. The response carries no analyzer version, so none is quoted. The 65.95% elsewhere on this page is a cost figure over analysed cost, a different denominator, so the two are not comparable. The trace is public, so every figure here can be reproduced.
how it works

Cheap filter first, expensive check second.

Boxdawn reads your trace as a graph and runs a two-stage cascade. The structural pass finds candidate waste in microseconds. The semantic pass looks only at what survived, never at everything.

Deterministic passFlow diagram. One trace file enters three deterministic detectors that run byte-exact, offline and in linear time; only the candidates that survive them are handed to an opt-in LLM judge, which is off by default; the result is a report.
pattern
repeat

The same agent runs on the same inputs again, with no new information in between. The cost breakdown in the JSON report keys it provable_duplicate, which is why the dashboard still joins on that name; the waste-rate block beside it uses repeat.

pattern
context_resend

Every API call re-sends the whole conversation. On long agent sessions this compounds to ~97% of input tokens as re-sent context (byte-exact chunk match).

pattern
redundant_read

The same file is Read multiple times without intervening writes. Common in coding agents that lose track of what's already in context.

pattern
duplicate_creation

A tool that changes something outside the agent runs twice on the same input, and the two runs come back with different ids: the second call created a second thing. A matching id is the same object, and a tool that returns no id is not judged either way.

pattern
semantic_duplicate

Two message chunks with different bytes but the same meaning (paraphrased re-send). Detected by the opt-in LLM-judge layer: off by default, requires API key.

Opt-in only · requires ANTHROPIC_API_KEY

pingpong requires genuinely multi-agent traces; single-agent coding sessions and Toolathlon are not expected to produce it.

pingpong is implemented but is not one of the cards above: the report gives a pair that name only when the pingpong detector itself produced it, and the waste-rate metric excludes it by name.

note: a fourth candidate pattern, regen_handoff, was descoped after failing pre-registered criteria. See Proven / Not Yet.


proven / not yet

What's proven. What's not.

Most tools hide this section. We think it is the product.

Proven
  • The problem is real. Token cost is a top operational pain in multi-agent systems, and the waste sits between agents, not inside them. (public market data, sources linked)
  • External evidence the problem exists. F1 0.72 on 1,575 traces from a single LangGraph stock-market application. Hybrid method, unrelated implementation (arXiv:2511.10650).
Not yet
  • Real-world semantic separation. In our real-data probes, precision came from the structural layer, where identity is a sha256 match; the semantic layer did not cleanly separate same-topic outputs. Against human labels precision is 0.826, and we do not claim every span it disagreed on is true waste. We published this finding instead of hiding it.
  • Measured savings. We have not yet measured a single dollar saved in production. No user count, no testimonials, because there are none yet.

When something moves from the right column to the left, you'll read it here first.

Waste tracks how long the agent works.

The share of the input bill that is resent context is not one number. It climbs with the length of the session, and it climbs without exception.

completed tool calls per session
  • 1-234.87%
  • 3-564.53%
  • 6-1077.67%
  • 11+88.02%

Measured on the fourth corpus only: 859 Claude Code sessions we did not collect, 2026-08-30. Only calls whose result came back are counted. The other three corpora have never been measured this way, so this is a curve inside one corpus, not a law across all of them.

open source

Open source, because trust is the product.

Boxdawn is MIT-licensed and deterministic by default. Four detectors (repeat, context resend, redundant read, and duplicate creation) run locally with no API calls and no signup. A fifth layer, semantic-duplicate LLM-judge, is opt-in only: off by default, requires ANTHROPIC_API_KEY and explicit CLEW_ENABLE_LLM_JUDGE=1 to activate. Same trace in, same report out. Every time, for the deterministic layers.

$ pip install boxdawn
$ boxdawn analyze your_trace.json

Works with OpenTelemetry SDK JSON and OpenInference trace exports: LangGraph, CrewAI, AutoGen, LlamaIndex. OTLP proto-JSON is not yet supported.

MITlocal-firstdeterministic-firstdeterministicopt-in LLM-judge1,000+ tests · CI greenOTel / OpenInference

about boxdawn

Why we started.

Servers were closed once. So were networks, and databases. Each got opened, and the industry that grew on top came after. Agents are the last layer still closed.

IBM Research wrote that “components such as agents, the LLMs powering them, and their associated tools often function as black boxes” in Formalizing Observability in Agentic AI Systems (AAAI 2026). Reading that left one question: it is a program, so why can we not see it?

What you cannot see, you cannot fix, cannot trust, and cannot hand work to. Enterprises are not withholding AI budget because they distrust it. They are withholding it because they cannot see it. Before AI can do a person's job, a person has to be able to watch it being done.

We start at cost for one reason: the bill is the only signal that leaks out of the black box.

Where we're going.

The goal is to make an agent hireable.

When you hire a person, you assume you can see what they did. You review it, you correct it, you pull back authority when trust runs short. When a person spends company money, a receipt is left and someone can be asked later what it was for. With an agent you can do none of that, and we keep handing agents more authority.

If AI is going to do a person's job, it has to answer the questions a person's job answers. The day it can, an agent stops being a tool and becomes staff.

How we work.

The discipline is pre-registration: criteria registered before results are seen, evaluations run exactly once, and negative findings published alongside the wins. The first version of this product failed its own pre-registered test. We killed it, said so publicly, and built this one on what we learned.

the ask

Running multi-agent in production? We want your trace.

Boxdawn is validated on 17,881 traces across four public benchmark corpora (28 Claude Code sessions, 6,780 Toolathlon trajectories, 10,056 Exgentic agent-LLM traces, and 1,017 Claude Code sessions we did not collect). The next honest step is private production traces from your team, and that's where you come in. Send us an execution trace (OTel or OpenInference JSON) and we'll run it through Boxdawn and send back exactly what it found, including false positives. No signup, nothing sold to you.

Can't share data externally? Fair. That's the point of local-first. Run Boxdawn yourself and share only the numbers. Either way, your trace directly shapes the recalibration, and we'll credit you publicly if you want.