Conversation
Supersedes PR KellerJordan#349. Based on record KellerJordan#89. Certified pool: 18 unseeded runs of the shipped source, every leg cold-cache, interleaved with 9 runs of record KellerJordan#89 on the same machine in the same session. this PR n=18 39.914 +/- 0.120 s val CE 3.27731 +/- 0.00099 record KellerJordan#89 n=9 73.889 +/- 0.137 s val CE 3.27828 +/- 0.00205 delta -33.98 s / -46.0 % / 1.85x One-sided t vs the 3.28 gate: t = 11.5, p = 9.4e-10 (17 dof). All runs counted. Logs and per-run statistics in records/track_1_short/2026-08-30_ANVIL2/.
devenpzak
changed the title
New Record: 0.665 minutes (39.9 seconds): ANVIL2, Sampled-softmax, New Embedding Table, Full-stack fp8
New Record: 0.665 minutes (39.9 seconds): ANVIL2, Sampled-softmax, New Embedding Table, Full-stack fp8 (-34.0s, -46% same hardware)
Aug 31, 2026
takuma104
added a commit
to takuma104/nanogpt-speedrun-rtx5090x1
that referenced
this pull request
Sep 3, 2026
…n#360) + 35 extension steps Shared-negative sampled softcapped cross entropy for training: per-micro-batch candidate set (all targets + stride-permutation negatives, P = 10240 / 10240 / 14336->24576 by stage, full softmax for the first 100 and the last 100 steps), lm_head fp8 rows gathered for the candidates, CE kernel compiled at VOCAB_SIZE=P, weight gradient densified. Validation is unchanged (full 50304-way softmax). 20 extra extension steps buy back the training-gradient bias. RTX 5090 x1: train_time 1212.1 s -> 1093.3 s (val_loss 3.2789). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JnTvksT3rz2Df7icStvcTu
… packed FP8 attention (KellerJordan#344) * perf(attention): pack reduced-width QK projections in FP8 Project Q/K at 96 dimensions while preserving 128-dimensional values. Fuse QK normalization, RoPE, layout conversion, and padding in Triton. Reuse dual FP8 layouts during backward to avoid extra transposes. Signed-off-by: Shivanjan Chakravorty <schakravorty846@gmail.com> * docs(track1): add H200 attention-packing evidence Signed-off-by: Shivanjan Chakravorty <schakravorty846@gmail.com> * perf(attention): finalize H100 track evidence Use 45 final-lr extension iterations to establish the required loss significance and replace the provisional H200 evidence with four fresh H100 runs. Signed-off-by: Shivanjan Chakravorty <schakravorty846@gmail.com> * perf(fp8): reduce training memory traffic Run all four MLP backward GEMMs through FP8 using cached row-major and transposed weight layouts. Refresh exact-current weight scales with Triton reductions and emit both layouts from a single weight read, while keeping activation scales lagged to avoid a mid-step synchronization. Eliminate the saved MLP pre-activation by reconstructing relu(pre) from the stored squared activation. Quantize logical QK/V packs without materializing concatenations or full-size abs temporaries, and reuse register-resident lane swaps in QK normalization and RoPE. Let the DC correction consume reduced-width Q/K views directly while retaining 128-wide V, avoiding padding copies and redundant score work. Signed-off-by: Shivanjan Chakravorty <schakravorty846@gmail.com> * docs(track1): add H100 record evidence with same-node baseline Seeds 2, 4, 42 and 1337 at 45 extension iterations give 3.274725 mean loss (p=0.001585) and 70.094 seconds (1.1682 minutes), against a 73.897 second (1.2316 minute) same-node f411b3d baseline. Make 45 the checked-in default so an unmodified run.sh reproduces the submitted 1315-step configuration. Signed-off-by: Shivanjan Chakravorty <schakravorty846@gmail.com> --------- Signed-off-by: Shivanjan Chakravorty <schakravorty846@gmail.com>
* Canonical token mask for val * Record 1xH100 runs * Timed * Drop extension steps 45 -> 40 --------- Co-authored-by: “ClassicLarry” <“larry36d@gmail.com”> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Conflicts with KellerJordan#344 and KellerJordan#350 resolved in favour of KellerJordan#360's trainer: train_gpt.py and triton_kernels.py are KellerJordan#360's as measured, dc_triton_kernels.py is removed (KellerJordan#360 retires the DC correction). README keeps records KellerJordan#90 and KellerJordan#91 and lists KellerJordan#360 as KellerJordan#92. KellerJordan#350's canonical masking is therefore not in the current trainer; its record and logs are unchanged.
ClassicLarry
added a commit
that referenced
this pull request
Sep 28, 2026
…373) Same trainer as the ANVIL2 record (#360), reorganized for readability. train_gpt.py is a linear main() (setup, model, warmup and CUDA-graph capture, timed loop, final validation); the model, optimizer, data and schedules live in track_1_short/, and the performance machinery (kernels, CUDA graphs, row prefetch, overlap, deferred gathers) lives under track_1_short/perf/, each file with a header on what it replaces and why it is faster. The run log embeds train_gpt.py and every file of the package. Tracks 2 and 3 are untouched. On the same 8xH100 nodes: #360 3.2769 / 40.60 s (n=4), this commit 3.2756 / 40.90 s (n=3). Intentional differences from #360: canonical token masking at the final validation (#350), dead MUDD gate lanes and value-embedding plane trimmed (init RNG stream differs), a claim-map n-gram gradient merge instead of a 16 GB dense buffer, the n-gram table kept outside the state_dict, TRAIN_SEED / NUM_SCHEDULED_ITERATIONS in place of KX_SEED / KX_STEPS, and fp8 / 8-GPU only. Co-authored-by: “ClassicLarry” <“larry36d@gmail.com”>
ClassicLarry
pushed a commit
to ClassicLarry/modded-nanogpt
that referenced
this pull request
Sep 28, 2026
Same trainer as the ANVIL2 record (KellerJordan#360), reorganized for readability. train_gpt.py is a linear main() (setup, model, warmup and CUDA-graph capture, timed loop, final validation); the model, optimizer, data and schedules live in track_1_short/, and the performance machinery (kernels, CUDA graphs, row prefetch, overlap, deferred gathers) lives under track_1_short/perf/, each file with a header on what it replaces and why it is faster. The run log embeds train_gpt.py and every file of the package. Tracks 2 and 3 are untouched. On the same 8xH100 nodes: KellerJordan#360 3.2769 / 40.60 s (n=4), this commit 3.2756 / 40.90 s (n=3). Intentional differences from KellerJordan#360: canonical token masking at the final validation (KellerJordan#350), dead MUDD gate lanes and value-embedding plane trimmed (init RNG stream differs), a claim-map n-gram gradient merge instead of a 16 GB dense buffer, the n-gram table kept outside the state_dict, TRAIN_SEED / NUM_SCHEDULED_ITERATIONS in place of KX_SEED / KX_STEPS, and fp8 / 8-GPU only.
ClassicLarry
pushed a commit
to ClassicLarry/modded-nanogpt
that referenced
this pull request
Sep 28, 2026
Same trainer as the ANVIL2 record (KellerJordan#360), reorganized for readability. train_gpt.py is a linear main() (setup, model, warmup and CUDA-graph capture, timed loop, final validation); the model, optimizer, data and schedules live in track_1_short/, and the performance machinery (kernels, CUDA graphs, row prefetch, overlap, deferred gathers) lives under track_1_short/perf/, each file with a header on what it replaces and why it is faster. The run log embeds train_gpt.py and every file of the package. Tracks 2 and 3 are untouched. On the same 8xH100 nodes: KellerJordan#360 3.2769 / 40.60 s (n=4), this commit 3.2756 / 40.90 s (n=3). Intentional differences from KellerJordan#360: canonical token masking at the final validation (KellerJordan#350), dead MUDD gate lanes and value-embedding plane trimmed (init RNG stream differs), a claim-map n-gram gradient merge instead of a 16 GB dense buffer, the n-gram table kept outside the state_dict, TRAIN_SEED / NUM_SCHEDULED_ITERATIONS in place of KX_SEED / KX_STEPS, and fp8 / 8-GPU only.
ClassicLarry
added a commit
that referenced
this pull request
Sep 28, 2026
Top section: under 40 seconds and under 330M tokens, the new techniques, and @devenpzak in the contributors list. Run instructions: the pinned cu128 torch, the CUDA 13 runtime the patched FA3 kernel needs, and the kernel-cache workaround for anonymous downloads. Record table: PR link and X handle for #92. Co-authored-by: “ClassicLarry” <“larry36d@gmail.com”>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.