
Amazon has expanded its family of foundational models with Nova Sonic, which has been developed to enable conversations with voices that are more similar to human ones based on speech comprehension and generation technologies.
As the company explained in its official blog, Nova Sonic is a new generative AI model with which Amazon seeks to simplify the development of voice applications. To this end, it offers a proposal that unifies comprehension and generation capabilities.
“This unification allows the model to adapt the generated voice response to the acoustic context and spoken input, resulting in more natural dialogue,” Amazon explained.
It can also understand nuances in conversations, including pauses and hesitations. In other words, the new model has arrived to make life easier for customers, as its primary purpose is to simplify the creation of voice applications, such as automated customer service calls and AI agents.
The system, integrated into Amazon Bedrock, is accessed through a new bidirectional streaming API and is intended for industries such as travel, education, healthcare, and entertainment.
Rohit Prasad, senior vice president of Artificial General Intelligence at the company, has reaffirmed the commitment to improving the user experience through voice-activated technology and stressed that the new model will make interactions more accurate, natural, and engaging.
In addition, the Nova Sonic from Amazon recognizes different speech styles. The company has emphasized that the AI model can also understand when a user speaks poorly, pauses, or mumbles. As of now, it only supports the English language. However, Amazon will soon add support for more languages.
The model has a context window of 32,000 tokens for audio, with an additional window to handle longer conversations.
Similarly, Nova Sonic demonstrates solid performance in overall conversation quality compared to other models in the industry, which at the moment include a select few with similar real-time conversational speech capabilities, such as OpenAI's GPT-4o (real-time) and Google Gemini Flash 2.0 (available through Gemini's experimental live API).
Companies such as ASAPP, Education First, and Stats Perform have already started integrating Nova Sonic to improve customer service, language learning, and sports data analysis, respectively. The companies have applauded the model's accuracy, low latency, and ease of integration.
On the other hand, Amazon has also launched the new Nova Reel 1.1 model that can now create longer videos based on text input.
The successor to last year's Nova Reel model, the new model can generate six-second shots, and a single video can have 20 such clips stitched together to create a 120-second video. It is also available to developers and general users through the Amazon Bedrock platform.
According to the company, this model improves creative productivity while helping to reduce the time and cost of video production through generative AI. It can be used to create engaging videos for marketing campaigns, product designs, and social media content with greater efficiency and creative control.
By continuing to use the site, you agree to the use of cookies. more information
The cookie settings on this website are set to "allow cookies" to give you the best browsing experience possible. If you continue to use this website without changing your cookie settings or you click "Accept" below then you are consenting to this.