NVIDIA Shows How to Make Large-Scale Model Pretraining Bitwise Deterministic at Low Cost
NVIDIA described a workflow built into Megatron Core that makes large-scale pretraining bitwise deterministic, meaning independent or checkpoint-resumed training runs follow the exact same numerical trajectory. Using tensor fingerprinting across training phases, modules, operations and kernels, engineers can trace nondeterminism down to a specific kernel and fix it. On a trillion-parameter Nemotron model running on 2,432 GPUs, optimizations brought the performance cost of enforcing determinism (the 'determinism tax') down to about 2%, while keeping bitwise-identical results over 800 training steps.