Research
Marginal tool utility in agentic debugging
Marginal tool utility and tool efficiency measure whether individual tool calls improve an agent’s probability of solving the task. Removing noisy tools preserved accuracy while doubling efficiency.
Root cause accuracy from observability data
A benchmark for the question every debugging agent should answer: what caused the production failure? Evaluated on root cause analysis from telemetry, not log summarization.
What debugging agents need from observability SDKs
A source-based comparison of Foam’s Node and Ruby packages with OpenTelemetry, Sentry, Datadog, New Relic, and Honeycomb.
- 01Debugging context
- 02Sensitive-data boundaries
- 03OpenTelemetry coexistence
- 04Instrumentation diagnostics
Clustering and noise reduction
How Drain3 streaming template mining compares to Sentry fingerprinting across 11 production environments.
Why AI SREs Make No Cents.
AI SREs need complete telemetry to reason well, but legacy observability pricing makes every additional byte of context more expensive.
Structural gaps in error monitoring: evidence from production systems
Evidence, causes, and a path forward for grouping, prioritization, configuration decay, alert noise, and AI-generated fixes.
- 01Clustering & duplicates
- 02Error prioritization
- 03Configuration decay
- 04Alert noise
- 05AI-generated fixes