Get the App
SLTechnology News&Howtos  ›  IT Information  › 

Microsoft creates a small LLM AI model with 1.3 billion parameters, which claims that the actual effect is better than that of GPT-3.5 with 100 billion parameters.

Shulou Source: shulou.com Published: 2023-11-24 17:43:38 10月04日 Update

CTOnews.com, June 27, AI model blind heap volume actually does not look better, more depends on the quality of training data. Microsoft recently released a 1.3 billion-parameter language model phi-1, which is trained with a "textbook-level" high-quality data set. It is said that "the actual effect is better than GPT 3.5 with hundreds of billions of parameters."

▲ Source ArxivCTOnews.com noted that the model was based on the Transformer architecture, and that the Microsoft team used "textbook level" data from the network and "logical content" processed with GPT-3.5, as well as eight Nvidia A100 GPU, to complete the training in just four days.

▲ source Arxiv Microsoft team said that rather than increasing the number of parameters of the model, it may be better to enhance the accuracy and efficiency of the model by improving the quality of the training data set of the model, so they used high-quality data to train the phi-1 model. In the test, the score of phi-1 is 50.6%, which is better than the GPT-3.5 (47%) with 175 billion parameters.

▲ image source Arxiv Microsoft said that phi-1 will next open source in HuggingFace, and this is not the first time Microsoft has developed a small LLM. Previously, they built a 13 billion-parameter Orca, trained using data synthesized by GPT-4, and performed also better than ChatGPT.

At present, the paper on phi-1 has been published in arXiv, and the relevant content of the paper can be found here.

Tags: Model training parameters Microsoft data reality effect content team education textbook grade paper quality rigor not necessarily next volume accuracy Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno Docker Microsoft OPPO Reno Linux Xiaomi