Nvidia: CPU is out of date, and the cost of training large language models with GPU can be reduced by 96%.
CTOnews.com, May 29 / PRNewswire-Asianet /-- according to Nvidia's speech at the 2023 Taipei computer Show, Nvidia claims that its GPU can significantly reduce the cost and energy consumption of training large language models (LLM).
Huang Renxun, CEO of Nvidia, challenged the CPU industry in his speech, believing that generative artificial intelligence and accelerated computing are the direction of future computing. He declared that the traditional Moore's Law is out of date and that future performance improvements will mainly come from generative artificial intelligence and methods based on accelerated computing.
Nvidia presented a total cost of ownership (Total Cost of Ownership,TCO) analysis of LLM at the show: first, they calculated the full cost of training a LLM server cluster of 960 CPU (including network, chassis, interconnect, etc.) and found that it cost about $10 million (CTOnews.com Note: about 70.7 million yuan currently) and consumes 11 gigawatt hours of power.
By contrast, if the cost remains the same, buying a $10 million GPU cluster can train 44 LLM at the same cost and less power consumption (3.2GWh). If you switch to keeping power consumption constant, you can achieve 150x acceleration through the GPU cluster, training 150 LLM at 11 gigawatt-hours of power consumption, but this costs $34 million, and the cluster occupies a much smaller footprint than the CPU cluster. Finally, if you only want to train a LLM, you only need a $400000 GPU server that consumes 0.13 gigawatt-hours of power.
What Nvidia is trying to say is that customers can train a LLM at 4 per cent cost and 1.2 per cent power consumption compared to CPU servers, which is a huge cost savings.