Skip to content
Friday, 9 October 2026
Tenesys AI News
Subscribe

NVIDIA Shows How to Make Large-Scale Model Pretraining Bitwise Deterministic at Low Cost

In short: NVIDIA described a workflow built into Megatron Core that makes large-scale pretraining bitwise deterministic, meaning independent or checkpoint-resumed training runs follow the exact same numerical trajectory. Using tensor fingerprinting across training phases, modules, operations and kernels, engineers can trace nondeterminism down to a specific kernel and fix it. On a trillion-parameter Nemotron model running on 2,432 GPUs, optimizations brought the performance cost of enforcing determinism (the 'determinism tax') down to about 2%, while keeping bitwise-identical results over 800 training steps.

Source: NVIDIA DeveloperNVIDIAOriginal article ↗

This summary was generated automatically by AI from NVIDIA Developer's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.

What changed?

  • 1Megatron-LM PR #7262 introduces ordered per-rank tensor fingerprinting to localize nondeterminism
  • 2Progressive tracing narrows divergence from training iteration down to phase, module, operation and kernel
  • 3Kernel fixes such as private output slots per grouped-GEMM writer plus fixed reduction order restore determinism without serializing all execution
  • 4Determinism tax reduced to approximately 2% at 2,432 GPUs with bitwise determinism maintained over 800 steps
  • 5Megatron Core adds recipe validation (--deterministic-mode), kernel testing and module validation across FP8/FP4 configurations to prevent regressions

Why it matters

Bitwise determinism lets engineers replay a loss spike exactly, validate that a system or software change had the intended effect, and resume interrupted training without silently altering results — all without paying a large performance penalty, which matters because at 10,000-GPU scale even a 10-percentage-point reduction in overhead can save on the order of 100,000 GPU-days.

What it means for AI agents and contact centers

This is relevant mainly to organizations training or fine-tuning large foundation models at scale rather than to teams consuming models via API for voice agents, so it has limited direct bearing on day-to-day voice AI or contact-center work.

Sources

  • NVIDIA DeveloperOfficialPrimary source
    „Scale Bitwise-Deterministic Pretraining with NVIDIA Megatron Core“
    6 Oct 2026, 22:58
    Original article →
Published by source
6 Oct 2026, 22:58
Found by our system
9 Oct 2026, 02:09
Summary generated
9 Oct 2026, 02:10

This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy

NVIDIA Shows How to Make Large-Scale Model Pretraining Bitwise Deterministic at Low Cost · TENESYS AI NEWS