Microsoft launches artificial intelligence model CoDi, which can interact and generate multimodal content
CTOnews.com, July 11, Microsoft recently released a press release called composable Diffusion Model (CoDi), a unique artificial intelligence model based on composable diffusion that is designed to interact and generate multimodal content.
Microsoft's goal of designing CoDi is to solve the limitations of the traditional single-mode AI model. In the case of synchronized video and audio, there may be inconsistency and alignment when independently generated information streams are spliced together.
CoDi adopts a unique combinable generation strategy to align multiple modes in the diffusion process to generate intertwined patterns. More importantly, CoDi can handle arbitrary input patterns and generate arbitrary modal content.
CoDi was developed by the Microsoft Azure Cognitive Services research team in collaboration with the University of North Carolina at Chapel Hill and is part of the Microsoft project i-Code, which uses artificial intelligence to enhance human-computer interaction.
CTOnews.com attached a link to the official introduction of the CoDi project, which can be read in depth by interested users.