Ordinary people can also become audio editors. Meta launched AI model Voicebox.
Thank you, Mr. Air, a netizen of CTOnews.com, for your clue delivery! CTOnews.com June 17 news, following the launch of ImageBind, Meta today launched a new generative AI model Voicebox. The model helps creators to perform voice generation tasks such as audio editing, sampling and stylization, which can be easily used by even ordinary users.
When introducing the Voicebox model, Meta said that the visually impaired can hear responses from friends, and ordinary users can speak a foreign language in their own tone and tone.
The AI model itself can generate high-quality audio clips, eliminate unnecessary background noise such as car speakers, retain the content and style of the audio, and use multiple languages to generate voice in six languages. Future developments of the model include providing natural sounds for visual assistants or non-player characters in meta-universe games.
Meta also compares Voicebox with other audio AI models such as Vall-E and YourTTS, which shows that Voicebox is more advanced and superior to the two models in comparing word error rate and style similarity.
CTOnews.com is here to attach a link to a detailed press release, which interested users can click to read.