A rounding slip in fast attention code quietly spoiled late training
Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training
TL;DR In a 450-million-parameter model, a fix called GProj cut the error a known repair left in one attention training signal from a median 219% to 0.34%, matching full precision. Trusted fast attention code can quietly derail training, and this fix adds 4.7% per training step.
The problem
Large AI models train in BF16, a compact number format that trades precision for speed. FlashAttention-3, a popular fast attention program, uses it. When the authors trained a model this way, all looked healthy until about halfway. Then the gradient, the signal telling training how to adjust the model, ballooned, and the model got worse. You would have seen no crash. A known repair stopped the blowup, but one gradient stayed badly wrong.
The idea
In attention, each word’s query is matched against earlier words’ keys. In exact math, a query’s gradient ignores where the keys sit as a group, because it combines them with weights that sum to zero. Picture surveyors combining teammates’ altitudes with weights summing to zero, so the mountain cancels and only height differences remain. Round the weights and altitude leaks in, more at higher camps. FlashAttention-3 rounds these weights, and late training pushes keys far out. GProj nudges the rounded weights back to a zero sum. That restored sum, the authors argue, removes a leak that grows as keys move from zero.
What they showed
- As accurate as full precision. On attention inputs captured from the failing run, the query gradient was off by a median 219% even after the known repair. GProj cut that to 0.34%, the level of full-precision attention.
- Training stays on track. In matched runs from scratch, GProj finished at the same training loss (prediction error, lower is better) as full-precision attention, while unmodified FlashAttention-3 ended 0.2 higher. The known repair alone, which GProj builds on, closed most of that gap.
- Cheap to run. GProj added 4.7% to each training step, against 33.4% for a fast full-precision version.
Why it matters
Teams pretraining with fast BF16 attention should check their gradients. On the same captured inputs, the authors found a similar error in other fast attention code, including PyTorch’s fused options, though they trained only with FlashAttention-3. No code is listed as released. For everyone else, it means a worse model can come from the math code itself and get blamed on the data instead.
The catch
The evidence comes from one small model on one GPU family. In a smaller standard model, attention gradients stayed close to full precision. The authors trace the failure to the very large queries and keys the main model grew, so runs that keep keys small may never hit it. The accuracy gaps come from inputs taken from the failing run; on inputs from healthy runs, every attention program was already accurate.
Who wrote it
First author Junlin Chen (Rutgers University and Carnegie Mellon University) wrote this with 11 colleagues from Rutgers University, Oracle, New York University and 2 other institutions. Chen also co-authored PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning (2026).