A coding AI improved by training only on its own post-mortems
Shockingly Simple Self-retrospection Improves Agentic Models Without RL
TL;DR On real Python repository issues, a 4-billion-parameter coding agent trained only to explain its own attempts solved 49.2% of problems, against 48.0% with a tuned reward-based method. Explanations may be useful training material, though the gap is small and reported without error bars.
The problem
Coding agents, AI models that fix software by editing files and running commands, are often sharpened with reinforcement learning. That means trying a task, getting graded by tests, and being nudged toward whatever passed. But a pass or fail doesn’t tell you which decision mattered. Sometimes it tells you nothing. GRPO, the version compared here, learns by comparing several attempts at one task, so if every attempt fails, the task teaches it nothing.
The idea
Can an agent improve by learning only to explain? Picture a cook who, after every dinner, burnt or perfect, writes a paragraph on which decision mattered, then leaves the notebook behind. Only the writing is practiced, but one cook does both jobs. In ROFT, the team’s method, the agent attempts a task, sees the test results, and writes a post-mortem: a key decision, the evidence, a correction. It is then fine-tuned (trained further) to predict only that text, with no reward. Later attempts see no notes. The authors argue that because the same internal settings produce explanations and actions, training one changes the other.
What they showed
- Slightly ahead of reinforcement learning. Qwen3.5-4B started at 44.2% on SWE-bench Verified, a set of human-checked Python repository issues. After training on problems picked to favor GRPO, ROFT edged out GRPO, as above, and also led on a harder benchmark.
- Learning from failure alone. On one bug in SymPy, a math library, where all 64 of the starting model’s attempts failed, training on that task raised its success rate from zero to 1.75%. GRPO gets no signal there. A Django run peaked near 1% and slipped back.
- Less training time. On identical hardware, ROFT finished the same number of training rounds in roughly half the time GRPO needed.
Why it matters
If you build coding agents, this is a training signal that, in single-task tests, still produced some successes when every starting attempt failed. The authors suggest it could reach past curated tasks with hand-built test checkers, turning ordinary interaction into training material. That step is untested here, but it would mean assistants that learn from routine work rather than only graded exercises.
The catch
The evidence is thin. ROFT’s lead on the main test is 1.2 points, reported without error bars. Each method is also scored at a different point in its training, so checkpoint choice could matter. Nothing checks whether a post-mortem is true, and the authors warn agents could absorb plausible but mistaken explanations. Experiments cover only coding, mostly with one small model.
Who wrote it
Jonathan Light, first and corresponding author, and nine colleagues from Microsoft Research, RPI, UC San Diego and two other institutions. Every author lists Microsoft Research. The paper is a preprint, and no code is public yet.