Note-sharing AI agents match far bigger solo crowds on puzzles

Scaling Discovery through Test-Time Communication

TL;DR On unfamiliar puzzle games, a team of five AI agents sharing notes fully solved games as often as the best of 33 agents working alone. It suggests agents that build on each other’s checked progress can beat more agents working apart, given enough compute.

The problem

AI agents are chatbot models that run code and tools on their own, for hours. Set several on a hard problem and the usual move is to run them separately and keep the best result. Each agent works blind, so a trick one finds early never reaches the others. Letting them talk sounds obvious, but sharing pulls against variety. Agents that swap notes may all pile onto the first idea that looks good.

The idea

Identical agents get one task and a shared folder: a message log, plus records of adopted ideas, failures and scores. Picture friends doing the same crossword in separate rooms. Solo, one friend must crack every clue; with a group chat, whoever fills a word posts it, crossing letters confirm it, and everyone builds on it. The prompt tells agents to copy a peer only after the task’s verifier, an automatic scorer, shows a clearly better result, and to keep one variation. Each team faces the best of equally many solo agents with the same resources. Once a score confirms a discovery, the authors argue, nobody has to rediscover it.

What they showed

  1. Puzzle games. On ARC-AGI-3, unfamiliar grid games with no instructions, five Claude Sonnet 4.6 agents fully solved games in 8.0% of tries, against 2.2% for the best of five solo agents. Matching them took the best of 33 solo agents.
  2. Beating the human best. Building the smallest digit-reading program that keeps 99.4% accuracy on the MNIST handwriting test, a GPT-5.6 Sol team reached 1,957 bytes. No solo run beat the best known human design, which the authors compressed to 2,461 bytes.
  3. Limits. Solo agents led at small budgets, a cost the authors call a coordination tax. In a small test on command-line tasks with unreliable feedback, a pair of agents beat one attempt but not two independent ones.

Why it matters

Anyone running AI agents on long optimization jobs with a clear score, like the packing-algorithm and model-shrinking tasks here, gets a simple setup worth trying: a shared folder and one prompt, which the paper prints in full. The authors’ hope is that parallel attempts become cumulative progress, the way researchers build on each other.

The catch

Several headline wins rest on few runs: the compression win is one team run against four solo runs, so a lucky run can’t be ruled out. The puzzle gains come from a handful of games, and most stayed unsolved either way. Teams also need enough compute and a clear score, and the authors leave open whether sharing helps when feedback is sparse, noisy or subjective. The puzzle and compression tests each used a single model.

Who wrote it

Jongho Park of UC Berkeley, working as a Microsoft Research intern, with four Microsoft Research colleagues: Vasilis Kontonis, Shivam Garg, Akshay Krishnamurthy and Dimitris Papailiopoulos. Park co-authored “Can Mamba Learn How to Learn?” (2024).