Meta developed the text generation image model CM3Leon, which is known as the best in the industry.
CTOnews.com July 16, Meta announced the development of an artificial intelligence model called CM3Leon, which can generate high-quality images based on text, generate text descriptions for images, and even edit images according to text instructions.
CTOnews.com Note: CM3Leon generation results (top) compare with DALL-E 2 generation results (bottom) Meta said that this model has reached the highest level in the industry in terms of text to image generation, surpassing the products of Google, Microsoft and other companies. CM3Leon is a model based on Transformer, and Transformer is a neural network structure that uses attention mechanism to process input data. Compared with other diffusion-based models, Transformer model is more efficient, faster training speed and lower computational cost.
Meta demonstrated CM3Leon's excellent performance in different tasks, including generating images based on complex text prompts, editing images according to text instructions, and generating image descriptions and responses. Meta said CM3Leon was a big step forward in image generation and understanding, but acknowledged that the model might have data biases and called on the industry to strengthen transparency and regulation.
Meta uses millions of licensed images from Shutterstock to train CM3Leon, and the most powerful version has 7 billion parameters, twice as many as OpenAI's DALL-E 2 model.
Meta did not say whether it would publicly release the CM3Leon model.