Get the App
SLTechnology News&Howtos  ›  IT Information  › 

GPT-4 model architecture disclosure: including 1.8 trillion parameters, using hybrid expert model

Shulou Source: shulou.com Published: 2023-11-24 18:04:03 10月03日 Update

CTOnews.com July 13 news, foreign media Semianalysis recently revealed the GPT-4 model released by OpenAI in March this year, including GPT-4 model architecture, training and reasoning infrastructure, number of parameters, training data set, token number, cost, hybrid expert model (Mixture of Experts) and other specific parameters and information.

▲ image source Semianalysis foreign media said that GPT-4 contains a total of 1.8 trillion parameters in layer 120, while GPT-3 has only about 175 billion parameters. In order to maintain a reasonable cost, OpenAI uses a hybrid expert model to build.

CTOnews.com Note: hybrid expert Model (Mixture of Experts) is a kind of neural network, which trains multiple models separately according to the data. After each model is output, the system integrates these models into a single task.

▲ map source Semianalysis it is reported that GPT-4 uses 16 hybrid expert models (mixture of experts), each with 111 billion parameters, and each forward route passes through two expert models.

In addition, it has 55 billion shared attention parameters and is trained with a dataset containing 13 trillion tokens. Tokens is not unique and is calculated as more tokens based on the number of iterations.

The context length of the GPT-4 pre-training stage is 8kPower32k, which is the result of fine-tuning 8k, and the training cost is quite high. Foreign media said that 8x H100 can not provide the required dense parameter model at the speed of 33.33 Token per second, so training this model needs to lead to a very high reasoning cost. If the H100 physical machine is calculated at US $1 per hour, then the training cost will be as high as US $63 million (about RMB 451 million).

In response, OpenAI chose to use the cloud-based A100 GPU training model, which reduced the final training cost to about $21.5 million (about 154 million yuan), and reduced the training cost in a slightly longer time.

Tags: Models training parameters costs experts mixing data people RMB systems reasoning output architecture upper and lower context two cloud tasks information including Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno vpn Docker Linux Microsoft MariaDB