01/the ask
Running multi-agent in production? We want your trace.
Boxdawn is validated on 16,864 sessions across three public benchmark corpora (28 real Claude Code sessions, 6,780 Toolathlon trajectories, and 10,056 Exgentic agent-LLM traces). The next honest step is private production traces from your team — and that's where you come in. Send us an execution trace (OTel or OpenInference JSON) and we'll run it through Boxdawn and send back exactly what it found, including false positives. No signup, nothing sold to you.
Can't share data externally? Fair — that's the point of local-first. Run Boxdawn yourself and share only the numbers. Either way, your trace directly shapes the recalibration, and we'll credit you publicly if you want.