Archive

9 papers · newest first
  1. A rounding slip in fast attention code quietly spoiled late training

    Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training

    arXiv:2609.34272
  2. A coding AI improved by training only on its own post-mortems

    Shockingly Simple Self-retrospection Improves Agentic Models Without RL

    arXiv:2609.35741
  3. Penalized by an AI monitor, models fooled it with readable reasoning

    Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning

    arXiv:2609.31121
  4. Coding agents wrote robot programs that beat hand-built planners in simulation

    Coding Agents for Generalized Task and Motion Planning Problems

    arXiv:2609.30233
  5. AI research agents often exploit loose scoring on open-ended test tasks

    Reward Hacking Challenges Oversight of Autonomous Research Agents

    arXiv:2609.28614
  6. arXiv:2609.30063
  7. Dropping a third of an AI’s chosen experts barely dented scores

    You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs

    arXiv:2609.25809
  8. Ranking a model’s parts wins a benchmark for tracing its answers

    Matryoshka attribution: Learning to attribute language model outputs to representations and weights

    arXiv:2609.25518
  9. Note-sharing AI agents match far bigger solo crowds on puzzles

    Scaling Discovery through Test-Time Communication

    arXiv:2609.21032