Get the App
SLTechnology News&Howtos  ›  IT Information  › 

Nesting dolls are not desirable: researchers prove that the results generated by practical AI training AI will lead to model degradation and even collapse

Shulou Source: shulou.com Published: 2023-11-24 17:30:08 10月03日 Update

CTOnews.com June 14 news, CTOnews.com friends may have imagined that if you use the results generated by AI to train AI, "doll training", what kind of results can be obtained? At present, a research team has actually observed and recorded this, and the detailed paper and the results have been published on arXiv.

In a word, "using the content generated by the model in training will lead to irreversible defects in the subsequent generated model." in human language, the researchers found that "training AI with the results generated by AI will only make the model worse and worse."

▲ image source arXiv it is reported that researchers have specifically studied the probability distribution of AI generation models, mainly around "text-to-text" and "image-to-image", and finally come to the conclusion: "because the results generated by each model have certain characteristics, the model generated by AI is used to train AI, and the latter will forget the real underlying data distribution over time."

Ilia Shumailov, one of the lead authors of the ▲ source arXiv paper, also said that "over time, errors in generating data (such as false examples) will force AI to further misperceive reality, and we are surprised to observe that model crashes occur so quickly that models can quickly forget most of the raw data they originally learned."

However, friends may wonder whether the model "degradation" can be avoided if the results generated by AI are artificially retouched and then put into model training.

The answer is no, the researchers found that "the process of model degradation is inevitable", so even for "idealized AI output after retouching", the model will be degraded after long-term learning.

For any large model, because there are too many learning data, they will inevitably be exposed to other AI-generated data, so the researchers said that "AI identification should be introduced to pick out the learning data that may be wrong" to improve the learning ability and accuracy of the model.

Tags: Model generation result research training data learning personnel researcher error inevitable content image partner text time paper passage observation Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno NVidia OPPO Reno Microsoft Shulou Technology Shulou Tech Info