# Cutting Pretraining Costs by 62% Introducing ANVIL III, our state-of-the-art LLM optimizer. September 3, 2026 Authors: Deven Pietrzak, Zafar Mamat, AJ, Mason Eyler, Nahom Seyoum, Max Misterka, Adit Srivastava This is Part 1 of our pretraining series, focused on the optimizer. We’ll cover our architecture changes in a future post. ## TL;DR - We cut pretraining costs by 62% at frontier scale (8x Chinchilla [1]), and we further reduce all AI inference costs by 30% (Faster decode). - ANVIL III marks the biggest optimizer breakthrough since Muon in 2024 and beats a fully tuned Muon by 20–28 millinats, regime dependent, at the same training budget. - Our model Feather 1.7B beats Qwen3-1.7B-Base on math, using 180x fewer total training tokens. - The compute efficiency from ANVIL III alone scales at a fixed 8x Chinchilla budget: 32% at 124M, 39% at 350M, 39% at 774M, and 50% at 1.2B parameters. - With ANVIL II, we record the largest pretraining efficiency jump on the NanoGPT speedrun. Our record is a greater percentage leap than the past 45 world records combined. ## The Muon Lineage ANVIL III is our proprietary optimizer; the earlier public ANVIL II NanoGPT speedrun submission is a separate version. The headline results combine ANVIL III with our internal architecture changes and are reported separately from the optimizer-only comparisons with Muon, whose fitted compute savings at a common 8× Chinchilla budget are 32%, 39%, 39% and 50% at 124M, 350M, 774M and 1.2B, respectively. The architecture changes produce a larger qualitative change, so we trained Feather using our combined methods and compare it with existing models. The Feather results therefore measure the combined optimizer and architecture. Muon comes from the matrix-preconditioned family of optimizers. An earlier optimizer in this family is Shampoo [2], which applies the update $L^{-1/4}GR^{-1/4}$. The original method accumulates gradient products; in the exponential moving average (EMA) variant, $L = \operatorname{EMA}(GG^T)$ and $R = \operatorname{EMA}(G^TG)$ [3]. If we write $G = U\Sigma V^T$ and remove this accumulation, we get $(U\Sigma^2U^T)^{-1/4}U\Sigma V^T(V\Sigma^2V^T)^{-1/4} = UV^T \tag{1}$ Equation (1) is the Muon update. Muon was introduced in 2024 by Jordan et al. [4]. The Muon update is the orthogonalization of the momentum update $\operatorname{EMA}(G)$. Since standard SVD algorithms are significantly slower and less parallelizable than matrix multiplication on GPUs, it is not practical to orthogonalize by SVD at each step [4]. For this reason, Muon orthogonalizes $\operatorname{EMA}(G)$ using Newton-Schulz iteration, which iteratively applies a polynomial to $\operatorname{EMA}(G)$. The polynomial is designed to send the matrix's singular values closer to $1$ at each iteration. Orthogonalization of the momentum updates is desirable because a large amount of useful gradient signal is contained in the middle singular components of the update matrix, which are traditionally orders of magnitude smaller than the highest singular components but which become on par with them after orthogonalization. Since the lowest singular components are mostly within noise of $0$, bringing their singular values up to $1$ is harmful, so the Newton-Schulz polynomial is only applied for five iterations, bounding the amount of inflation that these noise components receive. Attention stabilization+ In Kimi K2's experiments, attention-logit inflation occurred more frequently with Muon than AdamW. The report hypothesizes that growth in the spectral norms of the attention weight matrices $W_q$ and $W_k$ increases the risk of logit explosion. Standard approaches to prevent this inflation include adding QK-norm [5] to the architecture, which normalizes the $q$ and $k$ vectors before the softmax. However, this is an architectural change and does not work well with other architectural features such as multi-head latent attention [6]. An alternative to QK-norm is QK-clip, which scales each head's $W_q$ and $W_k$ matrices down whenever any of the softmax nonlinearities receive logits above a certain threshold in the current batch. Adding QK-clip to Muon gives MuonClip, the optimizer used to train Kimi K2 [6]. Muon uses a shared learning rate across the components of its matrix update. Rather than treating pretraining as a single optimization process in a high-dimensional loss landscape, Adam's adaptive update [7] in Equation (2) $\sum_{i, j} \frac{\operatorname{EMA}(e_i^\top G e_j)}{\sqrt{\operatorname{EMA}((e_i^\top G e_j)^2)} + \varepsilon} e_i e_j^\top\tag{2}$ treats pretraining as a Cartesian product of many smaller optimization processes and gives each process its own LR by normalizing its updates by their RMS across steps. NorMuon [8] introduces Adam's adaptive LR feature into Muon by dividing each row of the Muon update by the RMS of its L2 norm. There are two reasons why the orthogonalization doesn't automatically guarantee that this row norm is $1$. First, the Newton-Schulz variant called NS5, which is the standard choice of orthogonalization algorithm for Muon, has noticeable error, especially in directions with small singular values. Second, the row norms of orthogonal matrices are generally not equal to $1$ if the matrix is tall, and MLP layers have tall matrices. ## Beating Muon Across the Singular Spectrum ### The Adam SNR probe ANVIL III is an optimizer in the Muon family, where the final optimizer steps consist of orthogonal matrices. We pick paired orthonormal bases $\mathbf{u}_i, \mathbf{v}_i$ for each optimizer's update at a fixed checkpoint, before learning rate and weight decay, and measure them against the gradients of held-out batches, $C[b,i]=\mathbf{u}_i^\top\mathbf{G}_b\mathbf{v}_i$. An orthogonal matrix has every singular value equal to one, so $U$ and $V$ are only defined up to a shared rotation. We use three bases and report probes for all of them. - Momentum SVD: the singular vectors of the momentum before orthogonalization, the directions the momentum ranks highest. - Aligned basis: the rotation of that frame in which the training gradient is most diagonal. - Representative basis: a random rotation of the same frame. The update weights every direction in its frame equally, so this is what a typical direction of the step looks like. For the aligned basis, $Q$ is the eigenbasis of $S_X=\operatorname{sym}(U^\top G_X V)$, chosen on the training batch and fixed for held-out scoring. Descent inequality is the Gini coefficient [9] of $|d_i|$, where $d_i=u_i^\top G_Xv_i$ on the training batch, using the singular vectors of the momentum before orthogonalization. Lower Gini means more even contribution magnitudes across directions. Absolute values include ascent contributions. Using $C$, we define the Adam SNR of a singular vector in Equation (3). $\mathrm{adamsnr}_i = \frac{\operatorname{mean}_b C[b,i]}{\operatorname{rms}_b C[b,i]} \in [-1,1] \tag{3}$ A score near $1$ means the held-out batches agree on the direction. A score near $0$ means they do not. We also score random directions as a reference. Plot: Adam SNR of update directions. The distribution view plots probability density against mean(C)/RMS(C), from −1 to 1, over held-out batches. Rank views plot median directional SNR against within-matrix rank, ordered by the absolute training-gradient or momentum projection, with 16th–84th percentile bands. Controls select model size, training checkpoint, and pre-polar momentum SVD, aligned, or paired-random basis. Muon and ANVIL III are compared at matched size and token count. Dashed independent random-direction references are distinct from the paired-random basis. The accompanying Gini plot shows median Gini of absolute mean descent contributions versus training tokens, with 16th–84th percentile bands across matrices. It measures concentration in the stated basis, not the fraction of gradient energy captured. Figure 1. Adam SNR of the update directions on held-out batches in the selected basis, and descent inequality across update singular directions over training. ### The singular vectors We see that the Adam SNR distribution is persistently higher for ANVIL III and robust to the measurement methodology. When we rank by $|u^T G_X v|$, with $G_X$ being the gradient of the current step, the gap holds across the whole ranking and not just the largest projections. The size of the separation reduces across training, at the same speed that the Adam SNR reduces for random directions. Additionally, when we rank the singular vectors by their coefficients on the momentum (twin rail momentum for ANVIL III and both twin rail and standard momentum for Muon), we see that even the new singular vectors have a larger SNR, whereas new singular vectors found by Muon are low SNR. ### The weak singular directions Looking at the SNR plots, we wanted to know what the weakest directions were doing. We ranked directions by the magnitude of their momentum projections, removed the weakest $5\%, 10\%$, or $20\%$, and compared each optimizer with its own unmodified baseline. We see that ANVIL III is less sensitive to truncating the weakest singular vectors. kept | Muon | ANVIL III 95% | +1.1 | −1.2 90% | +2.5 | −0.1 80% | +6.9 | +1.8 Table 1. Change in validation loss (millinats) for runs trained from scratch with the weakest singular vectors removed by momentum alignment. ## Beating Muon at Every Tested Scale ### Experimental setup We measure the final validation loss of tuned Muon against ANVIL III. All runs use 200 warmup steps, linear decay to zero, a batch size of 524,288 tokens, and a sequence length of 1,024. Muon uses split QKV, which we find is 3–4 mnat better than fused QKV for Muon; ANVIL III uses fused QKV, so the comparison is conservative. ### Main results Plot: Final validation loss versus total training-token budget for Muon and ANVIL III at 124M, 350M, 774M, and 1.2B parameters. Lower loss is better. Points are measured final losses; lines connect measured budgets, not additional experiments. ANVIL III is 20 to 28 millinats better than Muon across the grid. To compare compute efficiency across model sizes, we use a common 8x Chinchilla budget. The fitted savings are 32%, 39%, 39% and 50% at 124M, 350M, 774M and 1.2B, respectively; the appendix gives the calculation. ### Training trajectories Plot: Validation-loss trajectories. The horizontal axis is consumed training tokens in billions and the vertical axis is validation loss in nats. Muon and ANVIL III curves retain recorded evaluations for the selected model size and training budget. Lower is better. Model and budget controls expose the available runs. Figure 2. Validation loss and the Muon − ANVIL III loss gap over training. Positive gap means ANVIL III has lower loss. The gap between ANVIL III and Muon is usually unpredictable for the first 20% of training. However, after the initialization and LR warmup, we see that the gap steadily increases from 30% to 80% of training, where, in certain configurations, the gap can grow as large as 50 mnat. Once the model is annealed, Muon partially catches up as it benefits more from annealing, as noisy methods do. ### Architecture Comparison, Compute matched We compare our architecture variant with the baseline at matched training compute, using four 712M models trained on 15.73B tokens. Optimizer | Baseline architecture | Architecture variant Muon | 2.9949 | 2.9655 ANVIL III | 2.9690 | 2.9393 Table 3. Validation cross-entropy (nats). The architecture improvement is nearly the same with either optimizer, and ANVIL III retains its advantage with both architectures. ## Beating Qwen3 as the Frontier 1.7B Math Model Using 180x Fewer Training Tokens. ## TL;DR - ANVIL III, coupled with architectural advances, allows us to plateau lower with 180x fewer total training tokens. - We compare Feather to SOTA open-source models on a bevy of math benchmarks including MATH500, GSM8K, GSM-Symbolic, Math-Perturb, SVAMP, and AMC23. - Feather does not regress on general reasoning capabilities. - We filter training documents against benchmark questions and solutions using lexical and number-normalized overlap rules. - Feather is an early preview of our architectures. ## Calibration - Our open-source data mixture is dramatically under-curated in contrast to Qwen's training corpus [10]. - In addition to distillation on model-generated data, Qwen is also aggressively midtrained on over 1T tokens of math data [12], while Feather has no midtraining stage and uses only 10B tokens of model-generated data. - Feather underwent three rounds of benchmark decontamination, covering every math benchmark in the comparison below. - Beyond Qwen2.5's documented overlap rules, our filters also target shorter question-and-solution matches and number-altered variants [11][12]. - Feather leads Qwen3-1.7B by 18.4 and 26.9 percentage points on GSM-Symbolic main and p1 with the available prompt settings. Feather also leads both similarly sized Qwen references on the two Math-Perturb variants with the available prompt settings. These results support generalization beyond the original benchmark questions. ## Introducing Feather 1.7B The comparisons below report Feather 1.7B (1.68B parameters, trained on 200B tokens) against Qwen3-1.7B-Base and Qwen2.5-Math-1.5B (trained on approximately 36T tokens [10] and up to 18T base-pretraining tokens [11] plus a math corpus of over 1T tokens [12], respectively). The math comparison uses the best available prompt setting for each model. We do this while preserving language and general reasoning capabilities. The language comparison uses zero-shot results, except for MMLU-STEM, which uses 5 shots. ### Full Qwen benchmark comparison Benchmark | n | Feather 1.7B | Qwen3-1.7B-Base | Qwen2.5-Math-1.5B | Qwen3-4B-Base MATH | 5000 | 60.8 | 52.6 | 52.4 | 52.4 MATH500 | 500 | 63.0 | 53.0 | 52.0 | 53.6 Perturb simple | 279 | 41.6 | 36.6 | 31.9 | 32.3 Perturb hard | 279 | 20.8 | 15.8 | 14.7 | 16.8 AMC23 | 40 | 32.5 | 22.5 | 27.5 | 35.0 AIME24 | 30 | 0.0 | 0.0 | 3.3 | 3.3 MMLU-STEM | 3153 | 36.3 | 61.8 | 46.4 | 75.3 AQuA-RAT | 254 | 23.6 | 31.1 | 26.0 | 41.7 SAT-Math | 220 | 25.5 | 40.0 | 33.6 | 45.9 MathQA | 2985 | 35.5 | 45.3 | 47.0 | 53.6 GSM8K | 1319 | 83.4 | 76.6 | 77.5 | 87.2 GSM8K Platinum | 1209 | 85.7 | 74.4 | 77.9 | 87.4 GSM-Plus mini | 2400 | 64.5 | 56.0 | 57.5 | 69.4 Symbolic main | 5000 | 81.1 | 62.7 | 69.5 | 82.5 Symbolic p1 | 5000 | 70.4 | 43.5 | 51.4 | 70.1 Symbolic p2 | 2500 | 40.1 | 20.2 | 29.0 | 50.3 SVAMP | 1000 | 87.4 | 74.3 | 85.1 | 87.6 MAWPS (integer subset) | 351 | 99.1 | 94.6 | 98.6 | 95.2 MGSM-en | 250 | 78.0 | 76.0 | 78.0 | 86.4 Feather 1.7B has 1.68B parameters and was trained on 200B tokens. Table cells below are plain accuracy values; formatting does not encode winners. Feather exceeds Qwen3-4B-Base on 6 of the 19 listed tasks: MATH, MATH500, Perturb simple, Perturb hard, Symbolic p1, MAWPS (integer subset). This is a task-specific comparison, not a win across the whole suite; MAWPS uses the integer-answer subset described below. These counts are descriptive, not significance tests. Feather scores refer to one 200B-token checkpoint. Generative results use common answer-scoring rules across all four models, and complete test-set counts were checked. Model validation checked the evaluation forward pass against the trainer validation loss. Multiple-choice tasks use fixed shot counts and scoring metrics. ### Language and general reasoning. Benchmark | Feather 1.7B(0.2T) | Qwen3-1.7B-Base(36T) | Qwen2.5-Math-1.5B(≤18T base + >1T math) HellaSwag · norm | 63.0 | 66.5 | 49.8 LAMBADA | 58.1 | 63.1 | 44.8 PIQA | 75.6 | 76.0 | 68.3 WinoGrande | 67.6 | 64.3 | 56.7 ARC-Challenge · norm | 46.6 | 44.9 | 40.7 OpenBookQA · norm | 40.8 | 38.4 | 34.8 SciQ | 95.8 | 96.1 | 92.0 MMLU-STEM · 5-shot | 36.3 | 61.8 | 46.4 “norm” = length-normalized accuracy; other rows use accuracy. Metrics are fixed across models. Feather scores use the complete archived task evaluations, with model validation performed separately using identical weights and evaluation implementation. ## Our Public NanoGPT Speedrun World Record A deprecated version of ANVIL directly contributes to our NanoGPT speedrun submission, and is entirely public. We see the NanoGPT speedrun as the most potent public indicator of pretraining efficiency for techniques used at the frontier. Our record submission beats the incumbent by 34 seconds on the same hardware: 73.889 seconds down to 39.914 seconds. Its 1.85× speedup is the largest single relative improvement compared with the published record history. By percentage reduction in training time, its 46.0% improvement is greater than the previous 45 world records combined: 45.8%, from record #44 to #89. Plot: NanoGPT speedrun record history. Bars compare training times and improvements across successive public records; selecting a record reveals its timing. The ANVIL II submission reduces the matched-hardware baseline from 73.889 seconds to 39.914 seconds. The figure compares this single improvement with the preceding 45 records combined; it is a whole-submission comparison, not an isolated optimizer ablation. Appendix: Tuning effort+ Muon was evaluated across learning-rate and weight-decay settings, with additional auxiliary Adam learning-rate checks. The recent large-model tuning used narrower muP proxies with the target depth and training-token budget; we varied individual hyperparameters around a reference setting and compared final validation losses. ANVIL III also had earlier learning-rate and weight-decay trials, and three completed 1.2B-target proxy runs compared auxiliary Adam learning rates. For the architecture comparison, hyperparameters were transferred from each optimizer's 774M/10B configuration in the main grid and used for both the baseline architecture and the architecture variant. The table counts completed comparison and tuning runs, including hyperparameter trials, their controls, and seed repeats. Component, architecture, data-order, and implementation ablation campaigns are excluded for both optimizers. These are run counts, not counts of distinct hyperparameter settings or an exhaustive history of optimizer development. Each trial is counted once across restarts. Smoke tests, timing-only tests, incomplete trials, and unrelated earlier optimizer implementations are excluded. Proxy runs are included under their target size; their losses are not full-size measurements. Each cell gives completed Muon / ANVIL III runs at the column's training-token budget, including muP proxies. Zero means no qualifying run in this audit. Target | 1.25B | 2.5B | 5B | 10B | 20B | 40B 124M | 37 / 12 | 46 / 9 | 23 / 8 | 38 / 8 | 16 / 6 | 5 / 1 350M | 0 / 0 | 0 / 0 | 10 / 5 | 37 / 9 | 10 / 12 | 5 / 1 774M | 0 / 0 | 0 / 0 | 0 / 0 | 15 / 1 | 11 / 2 | 8 / 2 1.2B | 0 / 0 | 0 / 0 | 0 / 0 | 20 / 1 | 12 / 2 | 7 / 4 Completed runs at the main token budgets: Muon / ANVIL III. Target | 1B | 2.62144B | 4B | 7B | 7.1429B | 14B | 15.5B | 28B 124M | 1 / 0 | 0 / 0 | 0 / 0 | 0 / 0 | 0 / 1 | 2 / 0 | 0 / 0 | 1 / 0 350M | 0 / 0 | 0 / 0 | 2 / 0 | 1 / 2 | 0 / 0 | 0 / 0 | 0 / 0 | 0 / 0 774M | 0 / 0 | 0 / 0 | 1 / 0 | 0 / 0 | 0 / 0 | 0 / 0 | 0 / 1 | 0 / 0 1.2B | 0 / 0 | 12 / 0 | 0 / 0 | 0 / 0 | 0 / 0 | 0 / 0 | 0 / 0 | 0 / 0 Completed runs at additional token budgets: Muon / ANVIL III. The 7.1429B column denotes 7,142,899,712 tokens; the 2.62144B column denotes 2,621,440,000 tokens. Appendix: Token-savings methodology+ At $1\times$ Chinchilla. Large open-weight models commonly train for several times this reference budget, so $1\times$ Chinchilla is less representative of their training regimes. We include it here as a comparison within our measured budget range. We set the baseline budget to $D_0=20N$ tokens and interpolate each optimizer's loss linearly in log token budget between adjacent completed runs. We first evaluate Muon's loss at $D_0$, then find the ANVIL III budget that matches that loss. Both budgets are bracketed by measured runs at every model size, so this comparison requires no extrapolation. Model | Muon budget (B tokens) | Matching ANVIL III budget (B) | Saving (%) 124M | 2.48 | 2.153 | 13.2 350M | 7.00 | 5.587 | 20.2 774M | 15.48 | 12.069 | 22.0 1.2B | 24.00 | 17.983 | 25.1 Optimizer-only compute savings at $1\times$ Chinchilla, interpolated within the measured budget range. The interpolation uses 1.25B–2.5B runs at 124M, 5B–10B at 350M, and 10B–20B at 774M. At 1.2B, it uses Muon's 20B–40B runs and ANVIL III's 10B–20B runs. These are interpolated equal-loss comparisons between separately annealed runs, not direct equal-loss stopping measurements. Savings are $100(1-D_{\mathrm{match}}/D_0)$, assuming equal compute per training token. At $8\times$ Chinchilla. We compare model sizes at a fixed $8\times$ Chinchilla budget, using $D_0=8\times20N$ tokens for a dense model with $N$ parameters. At each size, we fit $L(D)=C+AD^{-\beta}$ separately to Muon's and ANVIL III's completed 10B, 20B and 40B runs. Each run has its own learning-rate schedule, fully annealed at its endpoint. In the fit, $D$ is in billions of tokens. We carry forward the measured 40B optimizer loss gap, $\Delta=L_M(40)-L_A(40)$. For each fitted curve, the equal-loss budget is $D_{\mathrm{match}}=D_0\left(1+\frac{\Delta}{AD_0^{-\beta}}\right)^{-1/\beta}.\tag{1}$ We report the lower saving, $100(1-D_{\mathrm{match}}/D_0)$, from the two fits at each size. The fits supply alternative loss-curve slopes; the optimizer improvement is applied once. Model | Baseline budget (B tokens) | Loss gap (mnat) | Saving (%) 124M | 19.84 | 22.3 | 31.7 350M | 56.00 | 23.4 | 38.6 774M | 123.84 | 21.8 | 39.2 1.2B | 192.00 | 27.8 | 50.2 Optimizer-only fitted compute savings at a common $8\times$ Chinchilla budget. These are fitted equal-loss projections, assuming the loss-curve shape and optimizer gap persist over the relevant budgets. The 124M reference lies within the measured budget range; the larger models require extrapolation beyond 40B. Compute savings assume equal compute per training token. The measured final losses remain in the results table. Appendix: Architecture methodology+ #### Architecture methodology All four 712M models train for 30,000 steps on 15.73B tokens, approximately $1.1\times$ Chinchilla, with a sequence length of 512 and a batch size of 524,288 tokens. The runs use the same data order with seed 1. The setup differs from the main grid because these experiments were scheduled separately; those differences were incidental, not part of the architecture comparison. All models are evaluated on the same held-out text. Appendix: Headline numbers+ At 1.2B and $8\times$ Chinchilla, the fitted loss curves imply approximately 50% less training compute from ANVIL III. Combining this with the architecture's approximately 25% compute-saving estimate near $1\times$ Chinchilla gives approximately 62% less compute, assuming the gains compose. The fitted trends suggest larger optimizer savings at larger model sizes, but we do not include that potential benefit in this calculation. Appendix: Feather training mixture+ Feather was trained on 200B tokens in total. The final stage adds 22.985B tokens after 177.015B tokens of earlier training. The table describes the final stage's 25.019B-token input list, including deliberate repetitions; it is not a breakdown of the full 200B-token corpus. Training stops before exhausting that list, so listed shares are not exact per-source shares of the tokens consumed. Source category | Listed tokens (B) | Share Few-shot math packs | 2.965 | 11.9% Math instruction and reasoning datasets | 12.864 | 51.4% Benchmark training pools | 0.151 | 0.6% Language and STEM example renders | 3.892 | 15.6% Web and encyclopedia rewrites | 2.176 | 8.7% Books, stories and textbooks | 1.785 | 7.1% Other web and QA text | 1.186 | 4.7% Final-stage input-list composition. B denotes billion tokens. Math sources include OpenMathInstruct-2, NuminaMath, OpenMathReasoning, Nemotron-PrismMath, AoPS-Instruct, AceMath, DART-Math, MathScaleQA, MathCoder2, TemplateGSM, proof-pile and math web text, alongside examples rendered into multiple prompt formats. The language allocation includes benchmark training examples, STEM and QA material, synthetic rewrites, books and web text; it is not exclusively natural prose. The mixture includes teacher-generated solutions. For example, OpenMathInstruct-2 uses Llama-3.1-405B-Instruct generations. Selected sources are repeated two to four times. These are explicit list repetitions, not estimates of unique-document exposure across different renders of a source. Different renders can share underlying problems; the table measures listed tokens rather than unique examples. Appendix: Feather data decontamination+ Training-data preparation uses document-level benchmark-overlap filtering. The inherited corpus and final-stage sources have separate processing histories; the rules below describe the filtering methods, not a certification that every source is free of benchmark overlap. The original matching bank contains 103,101 items across 66 task/splits, covering all 19 math benchmarks in our comparison and language benchmarks. It includes test and validation data, selected development and few-shot rows, and questions, solutions, rationales, answer choices and support passages. We decoded training shards and scanned within document boundaries, checking hash matches against the actual normalized words. Matches removed whole documents. For packed few-shot examples, one flagged example removed the entire packed document. - `strict`: removed documents containing any matching 13-word sequence from a full benchmark item, at least 50% of a question's distinct 8-word sequences, or an exact short question with at least 30 characters and at most 7 words. - `v2strict`: added removal for matches covering at least 50% of a solution's distinct 8-word sequences, expanded benchmark coverage and recorded which field matched. - `v2lq`: normalized some LaTeX formatting and replaced numbers with 0, then removed qualifying long-question matches. Benchmarks commonly provide separate training and evaluation splits. The training splits are intended for model development; the held-out evaluation splits are used to measure performance. Feather's training included GSM8K training questions as exemplars and solution targets, MATH training solutions, MMLU auxiliary training data and other benchmark training pools. Using these designated training splits is standard practice. We also rendered examples in the evaluation prompt formats and generated multiple examples from some GSM training questions. Sharing a prompt format does not mean sharing the held-out questions or answers. Appendix: Benchmark methodology We evaluate base models without chat templates, using greedy decoding and BF16 forward passes. All four models use the same scoring rules. MATH answers are compared with the original dataset answers using symbolic equivalence and exact text matching after presentation cleanup. The last explicit answer takes precedence over earlier reasoning. Arithmetic answers use exact numeric equivalence, including fractions and decimals. For unanswerable GSM-Plus questions, explicit statements that the information is insufficient receive credit. Symbolic parse failures and timeouts count as incorrect. Multiple-choice tasks use answer likelihood, with length normalization where specified. Perturbation includes 279 examples per difficulty. We report the highest accuracy across the prompt settings below. A 4-shot prompt contains four worked examples; a zero-shot prompt contains none. - MATH, MATH500: Feather 0 / 4-shot, including training-style four-shot examples; Qwen 0 / 4-shot, Minerva prompt. - Perturb simple, Perturb hard, AMC23, AIME24: 0 / 4-shot, Minerva prompt. - MMLU-STEM: 5-shot likelihood. - AQuA-RAT, SAT-Math, MathQA: 0-shot length-normalized likelihood. - GSM8K: standard 0 / 5-shot; CoT 0 / 8-shot Feather also includes training-style few-shot examples. - GSM8K Platinum, GSM-Plus mini, Symbolic main, Symbolic p1, Symbolic p2, SVAMP, MAWPS (integer subset): 0 / 5-shot, GSM-style prompt Feather also includes training-style few-shot examples. - MGSM-en: 0 / 8-shot CoT. Full MATH accuracy pools all seven subjects within each prompt setting before taking the maximum. MAWPS uses the 351 integer-answer examples. ## References - Jordan Hoffmann et al. (2022). Training Compute-Optimal Large Language Models. - Vineet Gupta, Tomer Koren, and Yoram Singer (2018). Shampoo: Preconditioned Stochastic Tensor Optimization. - Jeremy Bernstein and Laker Newhouse (2024). Old Optimizer, New Norm: An Anthology. - Keller Jordan et al. (2024). Muon: An Optimizer for Hidden Layers in Neural Networks. - Alex Henry et al. (2020). Query-Key Normalization for Transformers. - Kimi Team (2025). Kimi K2: Open Agentic Intelligence. - Diederik P. Kingma and Jimmy Ba (2015). Adam: A Method for Stochastic Optimization. - Zichong Li et al. (2025). NorMuon: Making Muon more efficient and scalable. - Niall P. Hurley and Scott T. Rickard (2009). Comparing Measures of Sparsity. - Qwen Team (2025). Qwen3 Technical Report. - Qwen Team (2024). Qwen2.5 Technical Report. - An Yang et al. (2024). Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement. ### Comparing compute efficiency at fixed Chinchilla ANVIL III’s measured validation-loss advantage is approximately 20–28 millinats across the tested model sizes. We compare all model sizes at 8× Chinchilla, using D_0 = 160N tokens for these dense models. Taking the lower savings from the two fitted curves gives 31.7%, 38.6%, 39.2% and 50.2% at 124M, 350M, 774M and 1.2B, respectively. These are fitted projections; the raw validation losses are measured. The methodology appendix states the fitting procedure and extrapolation assumptions. ### Common Misconception: The 1.2B/40B Muon anchor is untuned or has an open LR/WD sweep The 1.2B/40B anchor is one completed full-size run, tuned with μP sweeps. Its matching five-setting LR/WD proxy sweep is complete, and the selected setting has the lowest measured proxy loss. The neighboring settings are 2.42–4.60 mnat worse; these differences are not a ±2 mnat uncertainty interval. The 350M seed check is a separate measurement, not the basis for declaring this LR/WD sweep complete. The loss-curve fits use fully annealed final-loss endpoints; the proxy sweep supports the Muon hyperparameter choice, not the extrapolation assumptions. ## Seed variation In our two-seed checks at 350M parameters and 5B and 20B training tokens, final validation loss differs by at most 1.1 millinats between seeds for either optimizer, compared with a 25–28 millinat gap between optimizers. The optimizer gap is much larger than the seed variation we measured. These are author-reported checks of those configurations, not uncertainty estimates for the entire model grid; per-seed losses are not included in the public plot exports. ## Numerical evidence and interpretation This appendix is generated from the public plot data used by the article. It includes every available model, budget, basis, and checkpoint, not only the initial interactive selection. ANVIL in the JSON series names means ANVIL III in the model/token-budget experiments; the historical ANVIL II speedrun submission is a distinct earlier version. Missing data are unmeasured or unavailable, never zero. ### Optimizer runtime overhead ANVIL III’s optimizer step adds less than 5% to the forward/backward runtime for the 1.2B configuration at a batch size of 524,288 tokens. The denominator is the forward/backward runtime. ### Final validation losses Loss is in nats/token; 1 millinat (mnat) = 0.001 nat. The difference below is computed from the displayed four-decimal table values. These cells are author-reported final results, not necessarily the endpoints of the separate paired trajectory runs. No per-cell seed uncertainty is supplied by this table. Model | Budget (B tokens) | Muon loss | ANVIL III loss | Muon minus ANVIL (mnat) --- | --- | --- | --- | --- 124M | 1.25 | 3.3876 | 3.3675 | 20.1 124M | 2.5 | 3.2694 | 3.2442 | 25.2 124M | 5 | 3.1803 | 3.1586 | 21.7 124M | 10 | 3.1202 | 3.0952 | 25.0 124M | 20 | 3.0733 | 3.0526 | 20.7 124M | 40 | 3.0483 | 3.026 | 22.3 350M | 5 | 2.9885 | 2.9627 | 25.8 350M | 10 | 2.9089 | 2.8825 | 26.4 350M | 20 | 2.8489 | 2.8225 | 26.4 350M | 40 | 2.807 | 2.7836 | 23.4 774M | 10 | 2.7923 | 2.7672 | 25.1 774M | 20 | 2.7227 | 2.698 | 24.7 774M | 40 | 2.6711 | 2.6493 | 21.8 1.2B | 10 | 2.7392 | 2.713 | 26.2 1.2B | 20 | 2.6615 | 2.6349 | 26.6 1.2B | 40 | 2.6059 | 2.5781 | 27.8 ### Complete machine-readable measurements The published JSON files are part of this article's evidence. Fetch the relevant file before making claims about individual curve points or alternate interactive selections. All stored numeric precision is retained. Hashes identify file contents, not independent experimental verification. This is the complete published numerical evidence, not raw gradients, checkpoints, or per-seed training logs. Dataset | File | Bytes | SHA-256 --- | --- | --- | --- charts | /evidence/anvil/charts.json | 150913 | f537df17120c3c4ae766395b2147c9839b71aed8a93543c06004d263f49a9e66 probes | /evidence/anvil/probes.json | 2293578 | 218cc9cc404bf0fcc6497023801f3e6b5020c011d05fffb30fd9f1b38aebe768 speedrun | /evidence/anvil/speedrun.json | 20524 | deecb2dfdf698ff4c41e89b84aa20603f2423c8f50ddbccb363b1a8f5f1b01bd Schema and measurement definitions: - charts.json: mainResults.series holds final-loss cells. figures contains validation-loss, validation-gap, and gini panels keyed by group (model label) and budget (billions of tokens). series.samples entries contain x, value, and optional low/high. Training-curve x is billions of consumed tokens. Loss values are nats/token; validation-gap values are 1000*(Muon loss - ANVIL loss), paired at exactly matching recorded tokens, without interpolation. Lines connect observations; they are not extra measurements. Empty scaling panels are not missing fitted experiments: no scaling-fit samples are published here. - probes.json: bases.pre_polar/aligned/random_polar.models lists model size, totalTokens, and checkpoints with exact tokens and progress (percent of training). A missing muon or anvil key means that optimizer was not measured at that checkpoint. For each optimizer, curves.hist and curves.null_hist store probability MASS per equal-width bin on [-1,1]; x is the bin center. The UI converts mass to density by dividing by 2/bin_count. Do not treat stored mass as the plotted density. curves.cx and curves.op contain within-matrix rank fraction x, median SNR value, and 16th/84th percentile low/high, pooled into 50 bins. Rank runs from largest absolute projection (0) to smallest (1). - Directional Adam SNR is mean_b(C[b,i])/RMS_b(C[b,i]), C[b,i]=u_i^T G_b v_i, using held-out gradients at fixed weights. Adam SNR is a metric, not the optimizer used to train these models. Momentum SVD uses the pre-orthogonalization momentum vectors. Aligned uses the eigenbasis of sym(U^T G_X V), selected on the training batch and fixed for held-out scoring. Representative basis applies a shared random rotation to the update frame. This is distinct from the independent random-direction references. nullStd is the standard deviation across reference directions; ±nullStd is not a confidence interval or a significance threshold. - Gini (figures.gini) uses |d_i|, d_i=u_i^T G_X v_i on the training batch, where u_i,v_i are the singular vectors of the pre-orthogonalization momentum operand. No singular-value weighting is applied. These are training-batch measurements, not held-out estimates. Training projections retain their stored precision; Gini is computed in float64. Each matrix contributes one Gini. value is the median across matrices; low/high are their 16th/84th percentiles, not uncertainty over training seeds. Absolute values include ascent. Muon uses separate Q/K/V matrices and ANVIL III fused QKV, so matrix populations differ. In our 124M/10B and 774M/10B Muon controls, split QKV improves final loss by 3–4 mnat over fused QKV with the other recorded settings matched. As noted in Experimental setup, using the stronger split-QKV Muon baseline makes the loss comparison conservative for ANVIL. The Gini matrix populations still differ. - speedrun.json: each record gives previousSeconds, seconds, multiplier=previousSeconds/seconds, date, public source link, retiming and submitted flags. These are whole-submission comparisons. The submitted 73.889-to-39.914-second improvement combines optimizer, architecture and systems changes; the submission estimates a 5.08-second optimizer-group attribution, including 0.92 seconds of graph capture, and separately reports 21 mnat at roughly equal wall clock. Its 164 ms/mnat conversion was selected to make component estimates sum to the total gain; these are not independent additive causal effects. The public plot data do not include a full ablation run log. The NanoGPT speedrun submission is tentatively approved by Keller Jordan. ### Muon tuning sweeps For the 1.2B/40B comparison, the published Muon result is one completed full-size run, tuned with μP sweeps, at LR 0.005 and WD 0.05, with final loss 2.6059268345. The matching μP LR/WD sweep is complete: all five settings were measured at 40B tokens on the width-512, 24-layer proxy, with auxiliary Adam LR held at 0.02. The selected LR/WD pair is best among those settings. Changing LR from 0.005 to 0.0035 or 0.007 increases loss by 4.04 or 2.42 mnat. Changing WD from 0.05 to 0.03 or 0.08 increases loss by 4.60 or 3.25 mnat. All four changes worsen loss. These are differences between hyperparameter settings, not plus/minus fluctuations at the chosen setting, seed variance, or an uncertainty interval for the full-size result. 'Sweep in progress' refers to other model/budget configurations, not this LR/WD sweep. The completed five-setting LR/WD sweep holds auxiliary Adam LR at 0.02. Separate auxiliary Adam-LR experiments are not included in this comparison. The 350M seed check measures seed variation at 350M; it is not used as an error bar for 1.2B. Likewise, the proxy sweep supports hyperparameter selection without establishing a numerical bound on full-size seed variation or proxy-transfer error. Proxy sweeps use narrower models with the target depth and token budget. Their losses are proxy losses, not full-size model results. Sweep-in-progress labels reflect unresolved tuning at roughly the 2-millinat level; they are not confidence bounds. Model or proxy | Budget (B) | Settings / status | LR | WD | Final loss (nats) | Completed runs --- | --- | --- | --- | --- | --- | --- 1.2B | 10 | Muon | 0.007 | 0.05 | 2.739211774338037 | 1 1.2B | 20 | Muon | 0.005 | 0.05 | 2.6631335452198983 | 1 1.2B | 20 | Muon | 0.007 | 0.02 | 2.66151369959116 | 1 1.2B | 20 | Muon | 0.01 | 0.05 | 2.66750682592392 | 1 1.2B | 40 | Muon | 0.005 | 0.05 | 2.605926834512502 | 1 124M | 1.25 | Muon | 0.014 | 0.05 | 3.399457210302353 | 2 124M | 1.25 | Muon | 0.014 | 0.1 | 3.393972 | 1 124M | 1.25 | Muon | 0.014 | 0.2 | 3.3910779267549516 | 2 124M | 1.25 | Muon | 0.02 | 0.05 | 3.394447 | 1 124M | 1.25 | Muon | 0.02 | 0.1 | 3.38763 | 3 124M | 1.25 | Muon | 0.02 | 0.2 | 3.396415 | 1 124M | 1.25 | Muon | 0.028 | 0.1 | 3.389619 | 1 124M | 1.25 | Muon | 0.028 | 0.2 | 3.4084248051047323 | 2 124M | 2.5 | Muon | 0.02 | 0.05 | 3.272103 | 1 124M | 2.5 | Muon | 0.028 | 0.02 | 3.278967 | 1 124M | 2.5 | Muon | 0.028 | 0.05 | 3.26959 | 3 124M | 2.5 | Muon | 0.028 | 0.1 | 3.279558 | 1 124M | 2.5 | Muon | 0.04 | 0.05 | 3.27575 | 1 124M | 5 | Muon | 0.0035 | 0.1 | 3.1856430530548097 | 1 124M | 5 | Muon | 0.0035 | 0.2 | 3.180290147662163 | 1 124M | 5 | Muon | 0.0035 | 0.4 | 3.18565194606781 | 1 124M | 5 | Muon | 0.005 | 0.1 | 3.184431 | 1 124M | 5 | Muon | 0.005 | 0.2 | 3.18150330632925 | 2 124M | 5 | Muon | 0.005 | 0.4 | 3.19230332672596 | 1 124M | 10 | Muon | 0.005 | 0.05 | 3.122532 | 1 124M | 10 | Muon | 0.007 | 0.02 | 3.129013 | 1 124M | 10 | Muon | 0.007 | 0.05 | 3.120193 | 3 124M | 10 | Muon | 0.007 | 0.1 | 3.120754 | 1 124M | 10 | Muon | 0.009 | 0.05 | 3.120832 | 1 124M | 20 | Muon | 0.004 | 0.05 | 3.076215 | 1 124M | 20 | Muon | 0.006 | 0.05 | 3.073252 | 2 124M | 20 | Muon | 0.006 | 0.1 | 3.07766 | 1 124M | 20 | Muon | 0.009 | 0.05 | 3.07515 | 1 124M | 40 | Muon | 0.007 | 0.05 | 3.048256 | 1 350M | 5 | Muon | 0.01 | 0.02 | 3.0051105678081513 | 2 350M | 5 | Muon | 0.01 | 0.05 | 2.992558 | 1 350M | 5 | Muon | 0.01 | 0.1 | 2.9929162546992303 | 2 350M | 5 | Muon | 0.014 | 0.02 | 2.996534 | 1 350M | 5 | Muon | 0.014 | 0.05 | 2.988465 | 2 350M | 5 | Muon | 0.014 | 0.1 | 2.995284 | 1 350M | 5 | Muon | 0.02 | 0.02 | 2.991506090760231 | 2 350M | 5 | Muon | 0.02 | 0.05 | 2.990621 | 1 350M | 5 | Muon | 0.02 | 0.1 | 3.006156726181507 | 2 350M | 10 | Muon | 0.01 | 0.02 | 2.914783 | 1 350M | 10 | Muon | 0.01 | 0.05 | 2.90896 | 3 350M | 10 | Muon | 0.014 | 0.02 | 2.9089055612683294 | 2 350M | 10 | Muon | 0.014 | 0.05 | 2.909628 | 1 350M | 20 | Muon | 0.006 | 0.02 | 2.855753757059574 | 1 350M | 20 | Muon | 0.006 | 0.05 | 2.8489235401153565 | 1 350M | 20 | Muon | 0.006 | 0.1 | 2.8550250053405763 | 1 350M | 40 | Muon | 0.0035 | 0.05 | 2.806994 | 1 774M | 10 | Muon | 0.007 | 0.05 | 2.793844 | 1 774M | 10 | Muon | 0.01 | 0.05 | 2.7923 | 2 774M | 10 | Muon | 0.014 | 0.05 | 2.795342 | 1 774M | 20 | Muon | 0.005 | 0.02 | 2.731248615682125 | 1 774M | 20 | Muon | 0.005 | 0.05 | 2.7227435171604157 | 1 774M | 20 | Muon | 0.005 | 0.1 | 2.7274400308728217 | 1 774M | 20 | Muon | 0.007 | 0.05 | 2.723155142366886 | 1 774M | 40 | Muon | 0.005 | 0.05 | 2.671134 | 1 774M | 40 | Muon | 0.007 | 0.05 | 2.675339 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.005 | 0.02 | 3.18678475022316 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.005 | 0.05 | 3.173629030585289 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.005 | 0.1 | 3.1653416752815247 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.01 | 0.02 | 3.1700692892074587 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.01 | 0.05 | 3.16146559715271 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.01 | 0.1 | 3.163354915380478 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.007 | 0.02 | 3.1784132957458495 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.007 | 0.05 | 3.16693320274353 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.007 | 0.1 | 3.1631874680519103 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.014 | 0.05 | 3.158847278356552 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.014 | 0.1 | 3.167195922136307 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.007 | 0.2 | 3.1703442871570586 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.01 | 0.2 | 3.1795843183994292 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.02 | 0.05 | 3.1610144436359406 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.014 | 0.03 | 3.1592316955327986 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.02 | 0.03 | 3.1582849234342576 | 1 1.2B-target μP proxy (width 512, 24 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.02 | 0.08 | 3.171355774998665 | 1 1.2B-target μP proxy (width 512, 24 layers) | 20 | Auxiliary Adam LR 0.02 | 0.007 | 0.05 | 3.115922886133194 | 1 1.2B-target μP proxy (width 512, 24 layers) | 20 | Auxiliary Adam LR 0.02 | 0.005 | 0.05 | 3.1197782129049303 | 1 1.2B-target μP proxy (width 512, 24 layers) | 20 | Auxiliary Adam LR 0.02 | 0.01 | 0.05 | 3.1137130916118623 | 1 1.2B-target μP proxy (width 512, 24 layers) | 20 | Auxiliary Adam LR 0.02 | 0.007 | 0.03 | 3.1196491569280624 | 1 1.2B-target μP proxy (width 512, 24 layers) | 20 | Auxiliary Adam LR 0.02 | 0.007 | 0.08 | 3.115542748570442 | 1 1.2B-target μP proxy (width 512, 24 layers) | 20 | Auxiliary Adam LR 0.02 | 0.014 | 0.05 | 3.1149869084358217 | 1 1.2B-target μP proxy (width 512, 24 layers) | 20 | Auxiliary Adam LR 0.02 | 0.01 | 0.03 | 3.1148780554533007 | 1 1.2B-target μP proxy (width 512, 24 layers) | 20 | Auxiliary Adam LR 0.02 | 0.01 | 0.08 | 3.1177545547485352 | 1 1.2B-target μP proxy (width 512, 24 layers) | 40 | Auxiliary Adam LR 0.02 | 0.0035 | 0.05 | 3.0892052441835403 | 1 1.2B-target μP proxy (width 512, 24 layers) | 40 | Auxiliary Adam LR 0.02 | 0.005 | 0.05 | 3.0851665318012236 | 1 1.2B-target μP proxy (width 512, 24 layers) | 40 | Auxiliary Adam LR 0.02 | 0.007 | 0.05 | 3.0875912100076675 | 1 1.2B-target μP proxy (width 512, 24 layers) | 40 | Auxiliary Adam LR 0.02 | 0.005 | 0.03 | 3.0897653490304946 | 1 1.2B-target μP proxy (width 512, 24 layers) | 40 | Auxiliary Adam LR 0.02 | 0.005 | 0.08 | 3.08841615319252 | 1 774M-target μP proxy (width 320, 36 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.01 | 0.05 | 3.3015567660331726 | 1 774M-target μP proxy (width 320, 36 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.007 | 0.05 | 3.3075292497873305 | 1 774M-target μP proxy (width 320, 36 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.014 | 0.05 | 3.2991103708744047 | 1 774M-target μP proxy (width 320, 36 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.01 | 0.03 | 3.3069484919309615 | 1 774M-target μP proxy (width 320, 36 layers) | 10 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.01 | 0.08 | 3.29834531545639 | 1 774M-target μP proxy (width 320, 36 layers) | 20 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.005 | 0.05 | 3.265473935008049 | 1 774M-target μP proxy (width 320, 36 layers) | 20 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.0035 | 0.05 | 3.2732989996671678 | 1 774M-target μP proxy (width 320, 36 layers) | 20 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.007 | 0.05 | 3.2605470389127733 | 1 774M-target μP proxy (width 320, 36 layers) | 20 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.005 | 0.03 | 3.2727865040302277 | 1 774M-target μP proxy (width 320, 36 layers) | 20 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.005 | 0.08 | 3.2616063207387924 | 1 774M-target μP proxy (width 320, 36 layers) | 40 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.005 | 0.05 | 3.236844855546951 | 1 774M-target μP proxy (width 320, 36 layers) | 40 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.0035 | 0.05 | 3.2461164861917498 | 1 774M-target μP proxy (width 320, 36 layers) | 40 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.007 | 0.05 | 3.2349483162164687 | 1 774M-target μP proxy (width 320, 36 layers) | 40 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.005 | 0.03 | 3.2425364524126055 | 1 774M-target μP proxy (width 320, 36 layers) | 40 | Auxiliary Adam LR 0.02 · Sweep in progress | 0.005 | 0.08 | 3.2354209810495376 | 1 ### Trajectory coverage The table lists the first and last recorded evaluation points for each trajectory. Model | Budget (B) | Optimizer | Evaluations | First tokens (B) | Last tokens (B) --- | --- | --- | --- | --- | --- 124M | 1.25 | Muon | 100 | 0.012582912 | 1.249902592 124M | 1.25 | ANVIL | 100 | 0.012582912 | 1.249902592 124M | 2.5 | Muon | 100 | 0.025165824 | 2.499805184 124M | 2.5 | ANVIL | 100 | 0.025165824 | 2.499805184 124M | 5 | Muon | 100 | 0.050331648 | 4.999610368 124M | 5 | ANVIL | 100 | 0.050331648 | 4.999610368 124M | 10 | Muon | 100 | 0.100139008 | 9.999745024 124M | 10 | ANVIL | 100 | 0.100139008 | 9.999745024 350M | 5 | Muon | 100 | 0.050331648 | 4.999610368 350M | 5 | ANVIL | 100 | 0.050331648 | 4.999610368 350M | 10 | Muon | 100 | 0.100139008 | 9.999745024 350M | 10 | ANVIL | 100 | 0.100139008 | 9.999745024 350M | 20 | Muon | 100 | 0.200278016 | 19.999490048 350M | 20 | ANVIL | 100 | 0.200278016 | 19.999490048 1.2B | 20 | Muon | 100 | 0.200278016 | 19.999490048 1.2B | 20 | ANVIL | 100 | 0.200278016 | 19.999490048 774M | 20 | Muon | 100 | 0.200278016 | 19.999490048 774M | 20 | ANVIL | 100 | 0.200278016 | 19.999490048 1.2B | 10 | Muon | 50 | 0.200278016 | 9.999745024 1.2B | 10 | ANVIL | 50 | 0.200278016 | 9.999745024 1.2B | 40 | Muon | 200 | 0.200278016 | 39.999504384 1.2B | 40 | ANVIL | 200 | 0.200278016 | 39.999504384 ### Probe coverage Tokens are exact counts here. M=Muon available; A=ANVIL III available. The paired count reports checkpoints where both were measured, not independent training repeats. Basis | Model | Total tokens | Checkpoints | Paired checkpoints | Checkpoint tokens: available optimizers --- | --- | --- | --- | --- | --- pre_polar | 124M | 2500000000 | 9 | 9 | 262144000:MA, 524288000:MA, 786432000:MA, 1048576000:MA, 1310720000:MA, 1572864000:MA, 1835008000:MA, 2097152000:MA, 2359296000:MA pre_polar | 350M | 20000000000 | 6 | 6 | 1048576000:MA, 2097152000:MA, 5242880000:MA, 10485760000:MA, 15728640000:MA, 19922944000:MA pre_polar | 774M | 20000000000 | 5 | 5 | 524288000:MA, 2097152000:MA, 5242880000:MA, 10485760000:MA, 19922944000:MA pre_polar | 1.2B | 20000000000 | 5 | 5 | 524288000:MA, 2097152000:MA, 5242880000:MA, 10485760000:MA, 19922944000:MA aligned | 124M | 2500000000 | 9 | 9 | 262144000:MA, 524288000:MA, 786432000:MA, 1048576000:MA, 1310720000:MA, 1572864000:MA, 1835008000:MA, 2097152000:MA, 2359296000:MA aligned | 350M | 20000000000 | 6 | 6 | 1048576000:MA, 2097152000:MA, 5242880000:MA, 10485760000:MA, 15728640000:MA, 19922944000:MA aligned | 774M | 20000000000 | 5 | 5 | 524288000:MA, 2097152000:MA, 5242880000:MA, 10485760000:MA, 19922944000:MA aligned | 1.2B | 20000000000 | 5 | 5 | 524288000:MA, 2097152000:MA, 5242880000:MA, 10485760000:MA, 19922944000:MA random_polar | 124M | 2500000000 | 9 | 9 | 262144000:MA, 524288000:MA, 786432000:MA, 1048576000:MA, 1310720000:MA, 1572864000:MA, 1835008000:MA, 2097152000:MA, 2359296000:MA random_polar | 350M | 20000000000 | 6 | 6 | 1048576000:MA, 2097152000:MA, 5242880000:MA, 10485760000:MA, 15728640000:MA, 19922944000:MA random_polar | 774M | 20000000000 | 5 | 5 | 524288000:MA, 2097152000:MA, 5242880000:MA, 10485760000:MA, 19922944000:MA random_polar | 1.2B | 20000000000 | 5 | 5 | 524288000:MA, 2097152000:MA, 5242880000:MA, 10485760000:MA, 19922944000:MA ### Gini measurements Basis | Batch | Model | Budget (B) | Optimizer | Tokens (B) | Median | 16th percentile | 84th percentile --- | --- | --- | --- | --- | --- | --- | --- | --- momentum SVD | train | 124M | 2.5 | Muon | 0.262144 | 0.749796167914058 | 0.5818150172027099 | 0.8464179403940916 momentum SVD | train | 124M | 2.5 | Muon | 0.524288 | 0.7234644862297619 | 0.5408703119117307 | 0.8191320518369288 momentum SVD | train | 124M | 2.5 | Muon | 0.786432 | 0.6956488098141306 | 0.507907638259541 | 0.7948877568575331 momentum SVD | train | 124M | 2.5 | Muon | 1.048576 | 0.6556191942963692 | 0.44941912212112894 | 0.7600314785067155 momentum SVD | train | 124M | 2.5 | Muon | 1.31072 | 0.6658479785861917 | 0.4417434084917719 | 0.7536280594565367 momentum SVD | train | 124M | 2.5 | Muon | 1.572864 | 0.6320712818864109 | 0.39826306101142867 | 0.7163634518822662 momentum SVD | train | 124M | 2.5 | Muon | 1.835008 | 0.6196614005308919 | 0.3951655132784429 | 0.7045390410477997 momentum SVD | train | 124M | 2.5 | Muon | 2.097152 | 0.5981519822470283 | 0.3616069784148387 | 0.6585906752617474 momentum SVD | train | 124M | 2.5 | Muon | 2.359296 | 0.5857573069329042 | 0.3466456329503996 | 0.6424675585998798 momentum SVD | train | 124M | 2.5 | ANVIL | 0.262144 | 0.6586414308538456 | 0.5036038015414895 | 0.7508723046866382 momentum SVD | train | 124M | 2.5 | ANVIL | 0.524288 | 0.5973114640707523 | 0.44099042580056536 | 0.7156238549073285 momentum SVD | train | 124M | 2.5 | ANVIL | 0.786432 | 0.6306081733268346 | 0.4556592342927339 | 0.7265755304777145 momentum SVD | train | 124M | 2.5 | ANVIL | 1.048576 | 0.5835293939485979 | 0.4253538359691614 | 0.6979118340784667 momentum SVD | train | 124M | 2.5 | ANVIL | 1.31072 | 0.5785333501391212 | 0.4133222715218053 | 0.6828477440212415 momentum SVD | train | 124M | 2.5 | ANVIL | 1.572864 | 0.560503950557542 | 0.40611414568914134 | 0.6739187924477213 momentum SVD | train | 124M | 2.5 | ANVIL | 1.835008 | 0.5303068436503826 | 0.3873170969378201 | 0.6344040165585092 momentum SVD | train | 124M | 2.5 | ANVIL | 2.097152 | 0.49727100856852846 | 0.3665259750815105 | 0.604310276853176 momentum SVD | train | 124M | 2.5 | ANVIL | 2.359296 | 0.48674639730200897 | 0.3570989539617038 | 0.5976998809811057 momentum SVD | train | 350M | 20 | Muon | 1.048576 | 0.6337155818051872 | 0.3915055811270317 | 0.7088032456081085 momentum SVD | train | 350M | 20 | Muon | 2.097152 | 0.6283815755601807 | 0.38761982355815405 | 0.7125054571279839 momentum SVD | train | 350M | 20 | Muon | 5.24288 | 0.6016429779609442 | 0.3476787945634256 | 0.6657535386719301 momentum SVD | train | 350M | 20 | Muon | 10.48576 | 0.5914104851215847 | 0.33607409010985057 | 0.6319603751403541 momentum SVD | train | 350M | 20 | Muon | 15.72864 | 0.5612139927894794 | 0.31322665426934426 | 0.5952719388401988 momentum SVD | train | 350M | 20 | Muon | 19.922944 | 0.5646393497927225 | 0.3130157308670331 | 0.6006979090518568 momentum SVD | train | 350M | 20 | ANVIL | 1.048576 | 0.5743876406647077 | 0.3854475954386055 | 0.6729081896023421 momentum SVD | train | 350M | 20 | ANVIL | 2.097152 | 0.5842426890421526 | 0.38623850923253444 | 0.6675456854902143 momentum SVD | train | 350M | 20 | ANVIL | 5.24288 | 0.5406672719059717 | 0.35873525647915455 | 0.626605691916784 momentum SVD | train | 350M | 20 | ANVIL | 10.48576 | 0.533712107875882 | 0.3442146042048824 | 0.6080938834189163 momentum SVD | train | 350M | 20 | ANVIL | 15.72864 | 0.48843726018043543 | 0.32521116437627756 | 0.5769276942870845 momentum SVD | train | 350M | 20 | ANVIL | 19.922944 | 0.4644105369455324 | 0.31435615501577635 | 0.5543117089886185 momentum SVD | train | 774M | 20 | Muon | 0.524288 | 0.6888602693831527 | 0.44945377606558684 | 0.7709380301574441 momentum SVD | train | 774M | 20 | Muon | 2.097152 | 0.6295151566035557 | 0.3714519620241856 | 0.7057981754333675 momentum SVD | train | 774M | 20 | Muon | 5.24288 | 0.5953955407852749 | 0.3309631856325514 | 0.6503204540752279 momentum SVD | train | 774M | 20 | Muon | 10.48576 | 0.5889128494571322 | 0.3298378894447694 | 0.6366604388635485 momentum SVD | train | 774M | 20 | Muon | 19.922944 | 0.5565430522780777 | 0.3040692927346427 | 0.5907600090638611 momentum SVD | train | 774M | 20 | ANVIL | 0.524288 | 0.5702388107890626 | 0.41019843991054794 | 0.6936082579662005 momentum SVD | train | 774M | 20 | ANVIL | 2.097152 | 0.5738373676615733 | 0.36240878752657596 | 0.646326161035393 momentum SVD | train | 774M | 20 | ANVIL | 5.24288 | 0.5285564287739459 | 0.344996714208201 | 0.6176570701987573 momentum SVD | train | 774M | 20 | ANVIL | 10.48576 | 0.5110507435746692 | 0.3276470014312595 | 0.5902482580483885 momentum SVD | train | 774M | 20 | ANVIL | 19.922944 | 0.48147936782872747 | 0.3152665813812237 | 0.5673969561280334 momentum SVD | train | 1.2B | 20 | Muon | 0.524288 | 0.7421418575504017 | 0.4935612003792899 | 0.8161788357331904 momentum SVD | train | 1.2B | 20 | Muon | 2.097152 | 0.6628156483586339 | 0.39259551736375614 | 0.7138206879474983 momentum SVD | train | 1.2B | 20 | Muon | 5.24288 | 0.6189402935479932 | 0.35193534573702645 | 0.6618823579707864 momentum SVD | train | 1.2B | 20 | Muon | 10.48576 | 0.6097619269628951 | 0.348859006708368 | 0.6433238658728768 momentum SVD | train | 1.2B | 20 | Muon | 19.922944 | 0.5694168663163709 | 0.3168941351058936 | 0.5998738333846307 momentum SVD | train | 1.2B | 20 | ANVIL | 0.524288 | 0.5905601921870948 | 0.44913279955802965 | 0.7320930552995203 momentum SVD | train | 1.2B | 20 | ANVIL | 2.097152 | 0.5811348528819948 | 0.40095033268924685 | 0.6843480105169352 momentum SVD | train | 1.2B | 20 | ANVIL | 5.24288 | 0.5437010250211285 | 0.36101236518142266 | 0.6536302713096664 momentum SVD | train | 1.2B | 20 | ANVIL | 10.48576 | 0.5252775957441608 | 0.3445540122843072 | 0.6214531399324376 momentum SVD | train | 1.2B | 20 | ANVIL | 19.922944 | 0.47932158074456166 | 0.32862232358643956 | 0.5805642080653347 ### Figure labels and explanations The following text accompanies the interactive figures and includes the current editorial wording. Label | Text --- | --- probe.title | Held-out Adam SNR of update directions probe.basis.pre_polar.label | Momentum SVD probe.basis.pre_polar.note | The singular vectors of the momentum before orthogonalization, the directions the momentum ranks highest. probe.basis.aligned.label | Aligned basis probe.basis.aligned.note | The rotation of the update frame in which the training gradient is most diagonal. probe.basis.random_polar.label | Representative basis probe.basis.random_polar.note | A random rotation of the same frame. The update weights every direction in its frame equally, so this is a typical direction of the step. probe.view.hist.label | Distribution probe.view.hist.note | Every direction of every matrix, weighted equally. probe.view.cx.label | Ranked by \|uᵀG_Xv\| probe.view.cx.note | Ranked within each matrix by the current training-batch gradient, 0 largest to 1 smallest, and pooled into 50 rank bins. The line is the median and the shading is the 16th to 84th percentile. probe.view.op.label | Ranked by \|uᵀMv\| probe.view.op.note | Ranked within each matrix by the momentum going into orthogonalization, 0 largest to 1 smallest, and pooled into 50 rank bins. The line is the median and the shading is the 16th to 84th percentile. probe.toggle | Show random-direction references probe.reference.hist | Dashed curves are random orthonormal directions scored on the same gradients. probe.reference.unavailable | Random-direction measurements are unavailable at this checkpoint. probe.reference.rank | one standard deviation across random directions. main.title | Final validation loss by model size and token budget gini.scale | 0 means even contributions and 1 means a few directions dominate. We use absolute values, so ascent counts too. gini.note | Lines are the median across matrices and shading is the 16th to 84th percentile. Muon uses separate Q, K, and V matrices and ANVIL III uses fused QKV. probe.chart.hist.metric | Probability density probe.chart.rank.metric | Median directional mean / RMS probe.chart.hist.axis | Adam SNR = mean(C) / RMS(C) probe.chart.rank.axis | Within-matrix rank fraction probe.chart.zero | Zero mean projection probe.chart.cx.title | Adam SNR · ranked by \|uᵀG_Xv\| probe.chart.op.title | Adam SNR · ranked by \|uᵀMv\| probe.chart.band | 16–84% of directions in rank bin gini.measurement.title | Per-matrix Gini of \|dᵢ\| history.title | Speedup, record by record — across the past 18 months history.subtitle | March 2025–August 2026 · NanoGPT speedrun history.metric | Speedup multiplier history.earlier | Earlier records history.submission | ANVIL II submission history.reduction | 46% less training time history.status | tentatively approved by Keller Jordan gini.train | G_X is the training-batch gradient. gini.momentum | uᵢ and vᵢ are the singular vectors of the momentum before orthogonalization, fixed across batches. main.sizes | Curve pairs, top to bottom: 124M, 350M, 774M, and 1.2B parameters. main.legend | Points are measured final validation losses; lower is better. ### Evidence limits The published plot exports do not include per-seed losses or selection receipts, so they do not independently verify the article's seed-aggregation procedure or establish training-run uncertainty. The reported optimizer-only compute savings compare a common 8× Chinchilla budget across sizes, using the fitted methodology above. They are equal-loss projections, not separately measured equal-loss runs; extrapolation assumes the fitted loss slopes and measured optimizer gaps persist. The 1.7B external-model benchmark comparisons use different architectures and training data, as stated in the article, and cannot isolate an optimizer effect. Their tables retain task metrics, shot counts and token budgets in the text above. Aggregate evaluation scores and document cross-entropies alone do not establish statistical significance or prove absence of contamination. A full independent reproduction would additionally require training recipes, data mixtures, checkpoints, evaluation harness revisions and per-seed/per-example results not supplied by these plot exports. ## Checkpoint validation losses Held-out loss at fractions of the token budget for the paired curve runs. | Parameters / Tokens | Optimizer | 25% | 50% | 75% | 100% | | --- | --- | --- | --- | --- | --- | | 124M/1.25B | Muon | 3.891 | 3.649 | 3.504 | 3.388 | | 124M/1.25B | ANVIL III | 3.905 | 3.645 | 3.481 | 3.367 | | 124M/2.5B | Muon | 3.675 | 3.508 | 3.384 | 3.270 | | 124M/2.5B | ANVIL III | 3.660 | 3.482 | 3.347 | 3.245 | | 124M/5B | Muon | 3.510 | 3.378 | 3.280 | 3.182 | | 124M/5B | ANVIL III | 3.519 | 3.373 | 3.257 | 3.159 | | 124M/10B | Muon | 3.384 | 3.283 | 3.204 | 3.121 | | 124M/10B | ANVIL III | 3.390 | 3.278 | 3.181 | 3.096 | | 350M/5B | Muon | 3.372 | 3.218 | 3.102 | 2.989 | | 350M/5B | ANVIL III | 3.360 | 3.196 | 3.068 | 2.963 | | 350M/10B | Muon | 3.244 | 3.125 | 3.023 | 2.909 | | 350M/10B | ANVIL III | 3.239 | 3.106 | 2.987 | 2.883 | | 350M/20B | Muon | 3.140 | 3.043 | 2.957 | 2.850 | | 350M/20B | ANVIL III | 3.124 | 3.013 | 2.913 | 2.822 | | 1.2B/20B | Muon | 3.013 | 2.900 | 2.795 | 2.662 | | 1.2B/20B | ANVIL III | 2.994 | 2.863 | 2.742 | 2.635 | ## Cite this article Deven Pietrzak, Zafar Mamat, AJ, Mason Eyler, Nahom Seyoum, Max Misterka, and Adit Srivastava. (September 3, 2026). Cutting Pretraining Costs by 62%: Introducing ANVIL III, our state-of-the-art LLM optimizer. Hyperstition.