SDSignal Desk

Scale Bitwise-Deterministic Pretraining with NVIDIA Megatron Core

Oct 6, 2026, 12:58 PM · NVIDIA Developer

Image: NVIDIA Developer

NVIDIA cut the cost of making giant training runs exactly reproducible to about 2 percent, which turns a debugging luxury into something labs can afford to leave switched on.

Why it matters

Training a frontier model across thousands of GPUs rarely produces the exact same numbers twice. Tiny differences in the order operations finish can change low-level bits, and those differences can snowball. That makes it hard to replay a failure, debug a sudden loss spike, or confirm that a resumed run picks up exactly where it left off.

In a new developer post, NVIDIA describes how Megatron Core supports bitwise-deterministic pretraining, where independent runs and checkpoint-resumed runs follow the identical numerical path. Working on Nemotron workloads, the team cut the performance overhead of determinism from double digits to roughly 2 percent at 2,432 GPUs, while holding bitwise determinism over 800 steps. For Nemotron 3 Ultra, overhead fell from about 17 percent to 1.5 percent at 3,072 GPUs.

From the desk

This is unglamorous infrastructure work, and we think it deserves more attention than it will get. Reproducibility is a basic scientific virtue that large-scale AI training has often had to give up for speed. When a run goes wrong, being able to replay it exactly and change one thing at a time is the difference between understanding a failure and guessing at it.

The cost argument is what makes this practical. NVIDIA's own illustration: on a hypothetical 100-day run across 10,000 GPUs, cutting determinism overhead from 15 percent to 5 percent saves 100,000 GPU-days. At around 2 percent, it becomes plausible to keep determinism on for production runs rather than reserving it for debugging. The kernel-level fix the post highlights, giving parallel writers private output slots and combining them in a fixed order, is a nice example of restoring determinism without simply serializing everything.

There is a broader benefit too. Exact reproducibility makes it easier to verify claims about how a model was trained and to catch corrupted checkpoints or hardware faults. As models grow more consequential, being able to show precisely what happened during training is part of accountability, not just engineering hygiene.

The limits are worth stating. NVIDIA is clear that determinism is validated within the same hardware and software environment; change GPU generations or library versions and results may differ. And these figures come from NVIDIA's own workloads.

Our read: a solid step that makes reproducible training cheaper. I'm watching whether other labs adopt deterministic modes as a default for their largest runs.

Context

The post outlines a debugging workflow proposed in Megatron-LM PR #7262 that records ordered per-rank tensor fingerprints and narrows a divergence from end-to-end metrics down to a specific kernel. Megatron-LM also offers a deterministic mode flag and byte-for-byte kernel tests to guard against regressions.

Who feels it

AI labs
Low-overhead determinism makes failure replay and exact checkpoint resumes viable on production-scale runs.
Infrastructure engineers
The tracing workflow offers a structured way to localize nondeterminism across thousands of GPUs.
Auditors and researchers
Reproducible training runs make claims about model training easier to verify.

What to watch

  1. Whether PR #7262 is merged into Megatron-LM
  2. Overhead figures reported by teams outside NVIDIA
  3. Expansion of deterministic coverage for new kernels and low-precision formats
  4. Adoption of deterministic modes in other training frameworks

Read the original

Continue at the source.

NVIDIA Developer

Companies: NVIDIA