Get the App
SLTechnology News&Howtos  ›  IT Information  › 

Microsoft announced Text To Speech Avatar AI tool: can make virtual 3D digital human, based on Azure platform

Shulou Source: shulou.com Published: 2023-11-24 22:04:08 10月03日 Update

CTOnews.com, November 16, Microsoft launched an AI tool called "Azure AI Speech text to speech (TTS) avatar" for Azure AI Speech at the Ignite conference, which claims to generate a lifelike human avatar (digital human). At present, this tool has been opened to the public preview trial.

Microsoft said that by using Azure AI Speech text to speech (TTS) avatar, users can build avatars based on "input text and say content", and combine with real-life photo training to build "interactive chatbots" based on real people, which can be used in corporate marketing, business or customer service scenarios.

It is reported that the Azure AI Speech text to speech (TTS) avatar mainly consists of three modules, namely, the text analyzer, the TTS sound synthesizer and the TTS virtual avatar synthesizer:

The text analyzer first analyzes the text entered by the user and produces a phoneme sequence (phoneme sequence). Then the TTS speech model in the TTS sound synthesizer predicts the acoustic characteristics of the user's input text, and then synthesizes the sound. Finally, the lip image of the character is predicted according to the above acoustic characteristics by the neural network sound synthesis model Avatar, and finally the virtual avatar image is formed.

Microsoft explained that the production of traditional avatars is time-consuming and labor-consuming, requires a dedicated shooting environment, and the post-editing process is quite expensive. At present, using Microsoft's latest Azure AI Speech text to speech (TTS) avatar service, after building a model for the first time, users can make a variety of product introductions, interactive videos and so on as long as they enter text. With Microsoft Azure OpenAI Service and neural network TTS functions, it can also present a more natural interactive experience.

CTOnews.com found that Microsoft claimed, for example, that users could use Azure AI Speech TTS avatar to mass-produce a variety of video content, such as corporate culture films, product introductions or CEO's digital avatars at conferences. Can also make virtual live digital human, chat robot, business robot, or online teaching AI teachers and so on.

Microsoft said that Azure AI Speech text to speech (TTS) avatar is now available to Azure subscribers in a variety of languages, and users can choose their desired roles from the preset avatar options or customize their own avatars.

If users want to customize their avatars, they need to upload a batch of video clips of people, and the Azure platform will process these videos online to generate avatars. The role itself is separate from the audio source. Users can choose the default audio source provided by the government, or they can upload their own training audio source.

Related reading: "launched in December, Microsoft released Personal Voice: achieve user-built AI audio in as little as 60 seconds"

Tags: Users Microsoft text production people voice video input digital content synthesizer machine robot model sound source analysis tools business product enterprise Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno Xiaomi NVidia MariaDB Apple Microsoft