After training GPT-3 in 3.9 minutes, Nvidia H100 broke six MLPerf records again.
CTOnews.com NVIDIA today released a press release stating that its H100 GPU set six new records in the MLPerf benchmark.
CTOnews.com reported in June that 3584 H100 GPU clusters completed a GPT-3-based large-scale benchmark in just 11 minutes.
The MLPerf LLM benchmark is based on OpenAI's GPT-3 model and contains 175 billion parameters.
Lambda Labs estimates that training such a large model requires approximately 3.14E23 FLOPS of computation.
Nvidia's latest Eos AI supercomputer, equipped with 10752 H100 Tensor Core GPUs and NVIDIA's Quantum-2 InfiniBand network, trained GPT-3 in just 3.9 minutes, a full seven minutes faster than June's test results.
Nvidia's other record-setting achievement in the post is the progress made in "system scaling," which has increased efficiency to 93% through various software optimizations.
Efficient scaling is very important in the industry because achieving high computing power requires more hardware resources, and if there is not enough software support, the efficiency of the system will be greatly affected.