↓ Skip to main content
Who Gets Credit When an Agent Recovers? A New arXiv Paper Asks
Daily Signal 3 min read

Who Gets Credit When an Agent Recovers? A New arXiv Paper Asks

An arXiv paper on causal evaluation of recovery in LLM agents asks whether your agent's reliability belongs to the model or to the harness around it.

When an LLM agent fails and then recovers, who deserves the credit: the model, or the harness wrapped around it? A new arXiv paper argues harnesses can lose that signal.

The paper is “When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents” (arXiv 2610.00372). I’m working from the title and listing, so I won’t quote its results. The question it asks is the one that matters, and the history behind it is easy to trace.

The tempo has changed. The field used to ask whether an agent finished the task. Now it asks how the agent got back on track after a failure, and what actually did the getting back.

March 2023: Reflexion shows an agent can reflect in words on a failed attempt and do better on the next one. Recovery becomes something you can design.

May 2024: SWE-agent shows that the interface around the model, the commands it gets and the feedback it sees, shapes how well it works. The harness becomes a first-class part of the system.

October 2026: This paper, whose arXiv ID places it in that month’s batch, asks how to measure recovery causally once the harness is doing so much of the work.

Here is the through-line. Recovery used to live inside the model’s reasoning. It now lives in a loop: retries, error messages rewritten for the model, truncated tool output, injected reminders, checkpoints.

A pass rate after all that is a blend. You cannot tell how much came from the model and how much from the scaffolding.

That is the mechanism behind “losing the signal.” Causal evaluation means intervening. Take the same failure, change one thing, and compare outcomes. Remove the retry. Strip the rewritten error. Swap the model and keep the harness. Whatever moves the result is what caused the recovery.

My read is blunt. A pass rate on an agent benchmark is not a model score. It is a model-plus-harness score, and teams that treat it as a model score will pick the wrong model. Worse, they will blame the model when a migration breaks something the harness was quietly covering for.

If you ship agents, the practical move is an ablation. Turn off one recovery mechanism at a time and log what happens. The Forge write-up on guardrails around a small model shows how much of an agent’s reliability can sit in the scaffolding. The full agentic SDLC piece covers the orchestration layer where those loops live.

Prediction: before the end of 2026, at least one widely used agent benchmark or eval framework will report recovery separately from final pass rate. If none does, this paper stayed an academic point and the blended number is still the industry’s scoreboard.

I send one of these a day, with the research that changes how you build. To get it, subscribe.