OpenAI seeks partners to generate datasets for training AI models
Thank you, Mr. Air, a netizen of CTOnews.com, for your clue delivery! CTOnews.com, November 10, OpenAI announced that it will work with organizations to generate public / private datasets for training AI models, and that the data partnership aims to "enable more organizations to help guide the future of AI" and "benefit from more useful models."
CTOnews.com learned from the blog that OpenAI said: "in order to eventually make AI more secure and benefit all mankind, we want the AI model to have an in-depth understanding of all topics, industries, cultures and languages, which requires as wide a training dataset as possible."
As part of the data partnership programme, OpenAI said it would collect "large-scale" data sets that "reflect human society" and are currently not easily accessible online. While the company plans to work across a variety of models, including images, audio and video, it specifically seeks data that "expresses human intentions" across different languages, themes and formats, such as long writing or dialogue.
OpenAI said it would work with organisations if necessary to use optical character recognition and automatic speech recognition tools to digitize training data and delete sensitive or personal information if necessary.
OpenAI wants to create two types of datasets: an open source dataset that anyone can use in AI model training, and a set of private datasets used to train proprietary AI models.
According to OpenAI, private sets are suitable for organizations that want to keep data private but want OpenAI's models to better understand their domain; so far, OpenAI has worked with the Icelandic government and Mi Mi eind ehf to improve GPT-4 's ability to speak Icelandic and to work with the Free Law Project to improve its model's understanding of legal documents.