Coding agents wrote robot programs that beat hand-built planners in simulation
Coding Agents for Generalized Task and Motion Planning Problems
TL;DR On 16 simulated robot tasks with an expert-built planner, programs from the best AI coding agent solved 95% of unseen test layouts on average, against the planner’s 47%. Much of the hand engineering in robot planning may be automatable, though only simulation was tested.
The problem
A simulated robot must fetch a block. The nearby one is hemmed in by obstacles, and a free one sits on a distant table. Which block to choose and how to grip it without collisions constrain each other, a tangle called task and motion planning. Standard planners solve each new layout from scratch with rules and motion routines experts hand-write per environment. Methods that reuse lessons across layouts still need heavy tailoring.
The idea
Could an off-the-shelf coding agent, an AI that writes and runs its own code, do that tailoring? Think of a consultant who tinkers with a test rig on a fixed fee, then leaves a procedure staff follow without calling back. Each agent got a task description, a simulator and a fixed budget of model usage, but no source code. Its one program was then frozen and run on new layouts with no AI involved. The agents were Claude Code running Opus 5 and Codex running GPT-5.6 Sol or GPT-6 Astra. Free experimentation, the authors argue, lets an agent spot what repeats across layouts, so the program needn’t plan from scratch.
What they showed
- Beats the hand-built planners. Astra’s programs beat the planner on 15 of the 16 tasks that have one. The weakest agent, Codex with Sol, beat it on only 9 of them.
- Far less computing per layout. On tasks with a planner and varied object counts, Astra’s programs used 0.5 seconds of computing per layout on average, against 29 seconds for the planners.
- Experimenting matters. LLMGenPlan, an earlier method, used Opus 5 with its extended thinking switched off, the same budget, and even read the source code, but only saw fixed error reports. It averaged 28% success, and the Opus agent given source code beat it on all but one task.
Why it matters
If you build robot planners for new setups, this is aimed at you: if it holds, the authors suggest, a coding agent could cut much of the rule-writing experts now do by hand. The team released all code and the full agent prompts. More broadly, the authors frame it as a test of whether skill at writing software carries over to physical reasoning, and in simulation it largely did.
The catch
Everything ran in simulation with perfect knowledge of every object’s position, so cameras, sensor noise and real hardware are untested. The models’ training data are undisclosed, so they may have seen the benchmark code; the authors counter with logs showing probing, testing and heavy revision. And some 3D tasks that require sweeping or pouring many small objects remain largely unsolved without source code.
Who wrote it
First author Matteo Merler (Fondazione Bruno Kessler) and corresponding author Tom Silver (Princeton University) wrote it with colleagues, including researchers at Carnegie Mellon University and the University of Cambridge. Silver co-authored KinDER (2026), the benchmark behind most tests here.