# Training NanoGPT in 39.9 Seconds Source: https://hyperstition.cc/training-nanogpt-in-39-9-seconds September 27, 2026 Authors: Deven Pietrzak ## TL;DR - We train the NanoGPT model to the target loss in 39.914 seconds on 8×H100, down from the previous record of 73.889 seconds. By percentage reduction in training time, this is a greater improvement than the [previous 45 world records combined](https://hyperstition.cc/cutting-pretraining-costs-by-62-percent#our-public-results-title). - We derive gains from ANVIL II, sampled softmax and a larger sparse n-gram model. - For reasons proprietary to Hyperstition, we redacted methods that scale to the frontier from our public record, apart from ANVIL II. Learn more about its extensions: [Cutting Pretraining Costs by 62%](https://hyperstition.cc/cutting-pretraining-costs-by-62-percent). ## Introduction The NanoGPT speedrun measures how fast NanoGPT (a 124M-parameter language model) can reach a fixed validation loss (3.28 on [FineWeb](https://huggingface.co/datasets/HuggingFaceFW/fineweb)) on 8 H100 GPUs. The record primarily comes from three changes. ANVIL II is an algorithmic improvement to the optimizer and achieves an improved validation loss compared to Muon. Sampled softmax reduces the number of vocab elements the model has to score in both forward and backward passes for most of the run and a much larger sparse n-gram table gives the model more capacity to use local token patterns. The code and measurements are in the [NanoGPT submission](https://github.com/KellerJordan/modded-nanogpt/pull/360). Plot: Complete instrumented validation curves for 22 ANVIL II runs and eight record #89 runs, by training step or recorded node-normalized time. Bands show observed minimum and maximum losses. The comparisons below show how each component contributes by ablating from the finished trainer, removing one mechanism at a time. Plot: Removal experiments priced in equivalent seconds. At 164 milliseconds per millinat, sampled softmax with prefix cross-entropy and a fused loss kernel contributes 8.41 seconds, the larger n-gram table 7.06, ANVIL with averaging and optimizer graphs 5.09, FP8 4.95, the smaller model with reduced attention width and depth 3.82, CUDA graphs 2.99, sparse embedding updates and loading 1.20, and layer mixing 0.49. To study how the changes contribute to the 33.975-second improvement, the sequence below adds each change one after another, with the same optimizer update rule, loader, and 1,194 steps held fixed, reporting the final validation loss and the time taken. We can see that some consequential changes might not change the time per step too much but have a significant loss advantage, reducing the number of steps required. Plot: Controlled reconstruction of selected changes. Training time in seconds / final validation loss: starting configuration 68.04 / 3.2993; table capacity 68.50 / 3.2635; sampled softmax 61.24 / 3.2664; FP8, width and depth 44.52 / 3.2962; CUDA graphs 40.83 / 3.2966; weight averaging 40.89 / 3.2784; sparse value embeddings 39.90 / 3.2780. ## ANVIL II ANVIL maintains two streams of momentum, fast and slow, and applies 6 polynomial maps to flatten the update's singular spectrum. This reduces the imbalance between strong and weak matrix directions, and agreement with slow momentum describes where cautious weight decay is applied. Alongside this, I was testing ways to improve the weights evaluated at the end. Keeping a running average of weights over steps can be better than the last iterate because the readout, matrix banks, and value embeddings update at different times, and their averages need to follow those update schedules. ### Weight averaging Late-stage weight averaging helps reach a lower validation loss with a small addition to the runtime. Rather than use the last weights, I combine them with averages collected near the end of training. I measured the contribution of the averaging by testing two configurations of the ANVIL II trainer with the optimizer update rule and training schedule held fixed. The averaging procedure brought the loss to 3.27840, while without it the loss was 3.29660, with a training time increase from 40.831 to 40.891 seconds. I maintain an exponential moving average (EMA) of the weights near the end of training. Each update moves the average toward the current weights, giving recent weights more influence than older ones. For one parameter group, with decay $0\leq\beta<1$, the update is $\bar\theta_t=\beta\bar\theta_{t-1}+(1-\beta)\theta_t.$ Near a local minimum, the loss can be approximated by a convex quadratic. If late training weights fluctuate within that region, averaging them reduces the loss associated with their spread. Under this approximation, the averaged weights have no greater loss than the weighted average loss of the individual checkpoints. For the readout, I blend 35% of the final weights with 65% of the average. ## Sampled softmax For every token in the corpus, the model scores each vocabulary element (the vocabulary has 50,304 elements), and roughly one third of the arithmetic during training is accounted for by these matrix multiplications. Most of the elements in the vocabulary don’t have much of an effect on the update, not only when looking at a single token but also an entire batch. I reduced these matmuls by keeping every target token in a local batch and adding a deduplicated set of non-target tokens. The model then computes the output head’s forward and backward passes only over these vocabulary columns. As a result, we get a different softmax normalization term. Let the candidate set be $C$, the vocabulary set be $V$, and $z_j$ be the score for vocabulary element $j$. The softmax normalizers are $Z=\sum_{j\in V}e^{z_j},\qquad Z_C=\sum_{j\in C}e^{z_j}.$ Since $p_j=e^{z_j}/Z$, we get the probability mass assigned to the candidate set as $r_C=\sum_{j\in C}p_j=Z_C/Z$. So restricting our attention to the candidate set results in this distribution: $\widetilde p_j=\begin{cases} p_j/r_C,&j\in C,\\ 0,&j\notin C. \end{cases}$ This reduces the compute required but also makes us optimize a different objective. Given that the candidate set includes the target token $y$, we can calculate the difference between the full and restricted loss gradients with respect to the logits $z$. For cross-entropy, each score gradient is its predicted probability minus one if it is the target. Writing the full loss as $\ell$ and the restricted loss as $\ell_C$, $\frac{\partial\ell}{\partial z_j}=p_j-\mathbf 1\{j=y\},\qquad \frac{\partial\ell_C}{\partial z_j}=\widetilde p_j-\mathbf 1\{j=y\}.$ For omitted elements of the vocabulary, we keep the gradient at zero in the sampled softmax variant. Thus, we can bound the L1 norm of the difference in gradients with respect to the logits: $\begin{aligned} \lVert\nabla_z\ell_C-\nabla_z\ell\rVert_1 &=\sum_{j\in C}(\widetilde p_j-p_j)+\sum_{j\notin C}p_j\\ &=(1-r_C)+(1-r_C)=2(1-r_C). \end{aligned}$ So the more total probability mass the model has given to the candidate set, the smaller the error between the restricted and full-softmax score gradients. This lets us reduce the compute per step, potentially at the cost of final validation quality. In these experiments, expanding the candidate set over training and ending with the full vocabulary substantially reduced the validation loss gap. Jumping directly to the full vocabulary from a smaller candidate set worked worse than expanding the set gradually. The final schedule uses 10,240 → 14,336 → 24,576 candidates, followed by all 50,304 entries for the final 87 steps. The test was whether the model could close the validation loss gap with the full-vocabulary alternative by the end of the training run as the sampling set kept expanding. ## The bigram and trigram table Record #89 already used a bigram table, but it had plenty of room for reducing the compute and memory required. It maintained a copy of the table on each GPU, and each update involved a full-table-sized gradient, even though the batch responsible for that update only affected a few rows of the table. Thus, increasing the size and adding a trigram table was infeasible for a speedup. I made each GPU own one eighth of the weights and request only the rows which need to be updated in the upcoming batches. These vectors are kept in a local cache. During backward passes, repeated row IDs are combined, and their updates are sent to the GPU where they reside. These improvements make it feasible to greatly increase the sizes and even include a trigram table. Each row’s optimizer state occupied 6,144 bytes of space. I reduced this to a second-moment scalar and a timestamp per row. The update can use the last update’s timestamp to calculate the amount of second-moment decay it missed. This reduces the per-row state to 8 bytes. Separately, removing the full table replica saved 12.2 GiB per GPU. With these changes, I could increase the table size and add trigram channels according to how much they helped the validation loss compared to how much time they added per step. The final configuration has 84.6 million rows across bigram and trigram channels, compared to the original 377,280. The finished recipe with the original table size reached around 3.313 in the same number of steps as this configuration reaches 3.278, with little added time. This is a speedup of 7–8 seconds, since that is how much longer the smaller-table configuration takes to reach a validation loss of 3.28. The original table had to store information about several unrelated sequences in the same row and had a lot of collisions. In a sample of 100 million training tokens, a typical row in the original bigram table was shared by 30 different token pairs. In the larger table, a typical row held just one pair or triple, so unrelated patterns no longer had to share the same learned vector as often. ## Final results The complete trainer reduces training time from 73.889 to 39.914 seconds, saving 33.975 seconds, or 46.0%. We measured this 1.85× speedup in a comparison of 18 ANVIL II runs and nine interleaved record #89 runs on the same machine, with all 18 ANVIL II runs below the 3.28 validation target. The controlled sequence shown earlier studies selected changes with the final optimizer update rule, schedule, and loader held fixed. Its 28.15-second reduction measures those changes within the reconstruction; the full result comes from the direct comparison with record #89. Compilation and warmup happen before timing begins. Training time includes the final averaging and synchronization, while validation happens afterward using the full vocabulary.