QRazy
questions worth discussing
📰 News Question / 💻 Technology
This post was published on the Qrazy.net blogs. Its author is not affiliated with the Qrazy.net editorial team.

OpenAI and Alibaba Unveil New Voice Models

"Discover new innovative voice models from OpenAI and Alibaba. Learn how these technologies are transforming interaction with artificial intelligence and enhancing the quality of voice communication."

QRazy's Key Findings

  • OpenAI has introduced models for accurate speech recognition, adapting them to the conversation context and reducing the error rate.
  • Alibaba has developed systems for natural voice communication, allowing users to interrupt responses without referring to text.
  • The new models differ in their tasks: OpenAI focuses on transcription while Alibaba emphasizes interactive dialogue and the use of external functions.
OpenAI and Alibaba Unveil New Voice Models
Media question about ancient symbols

OpenAI and Alibaba have both recently showcased new models for voice interaction, though they are tackling different tasks. OpenAI is focused on improving speech-to-text translation, while Alibaba is developing systems that facilitate real-time voice conversations.

New OpenAI Models for Speech Recognition

OpenAI has introduced two models: GPT-Live-Transcribe and GPT-Transcribe.

GPT-Live-Transcribe transcribes speech in real-time during a conversation. The audio is processed as it comes in, so the text appears almost immediately. This model is suitable for calls, online meetings, voice assistants, and other services where low latency is important.

GPT-Transcribe, on the other hand, is designed for pre-recorded content. For example, it can be used to transcribe interviews, lectures, podcasts, or phone calls. The model is also suitable for batch processing a large number of audio files.

Models Can Be Provided Context in Advance

Before starting the transcription, the user can inform the model about the content of the recording. In the description, the user can include names of people, company names, professional terms, keywords, and expected languages.

This will be useful in cases where rare surnames, product names, or specialized expressions are present in the conversation. Without these hints, the system may misinterpret an unfamiliar word for a more common one, leading to incorrect transcription.

GPT-Live-Transcribe also considers prior exchanges. The model does not treat each phrase in isolation but instead tries to follow the flow of the conversation. If someone continues a previously started thought, the system can use already transcribed text to understand the next response.

What OpenAI's Tests Showed

In the OpenAI Context Aware ASR test, additional information about the recording significantly influenced the results.

The semantic accuracy of GPT-Live-Transcribe improved from 38.5% to 44.6%. For GPT-Transcribe, the score rose from 41.6% to 45.2%.

In real audio recordings, the new models also made fewer errors than their predecessors. The error rate for GPT-Live-Transcribe was 9.60%. In comparison, GPT-Realtime-Whisper showed an error rate of 11.65%.

GPT-Transcribe finished the same test with a result of 8.98%, while Whisper-1 reached an error rate of 15.21%.

According to Artificial Analysis, another test showed the word error rate for GPT-Transcribe was 3.31%, which is significantly better than the previous OpenAI models.

Processing 1000 minutes of audio using GPT-Transcribe costs $4.50.

Alibaba Bets on Live Conversation

Alibaba introduced Qwen-Audio-3.0-Realtime Plus and Qwen-Audio-3.0-Realtime Flash. These models are primarily designed for voice communication. A person speaks to the system, and it responds with a voice without the need to constantly switch back to a text interface.

The conversation does not necessarily need to be strictly turn-based. The user can begin speaking while the model is still responding. The system is expected to detect the new remark, halt the current response, and switch to what the person has said.

This is an essential capability for voice assistants. In typical systems, users often have to wait for a long response to finish or repeat a question if the assistant did not notice the interruption.

Alibaba's models are also capable of accessing external functions. This allows the voice assistant not only to answer queries but also to perform actions in connected services. Furthermore, the ability to clone user voices has been announced.

Qwen Tops the Rankings

Qwen-Audio-3.0-Realtime Plus topped the Artificial Analysis Speech-to-Speech Index with a score of 84.1%.

OpenAI's GPT-Realtime-2.1 High took second place with a score of 79.1%.

Alibaba's solution also demonstrated better results in tests for speech reasoning, dialogue management, and agency tasks. In other words, the model performed well not only in casual conversation but also in scenarios where it needed to understand context and use additional tools.

Major Drawback — Long Pause Before Response

Despite its strengths, Qwen-Audio-3.0-Realtime Plus has a notable drawback. The model does not start responding immediately.

According to Artificial Analysis, the first audio output appears approximately four seconds after the user's remark. In a regular voice conversation, such a pause is quite noticeable.

In contrast, GPT-Realtime-2 High has a significantly shorter delay — about 1.14 seconds.

Companies are Currently Developing Different Directions

OpenAI is currently focusing on accurate speech transcription. The new models are suitable for processing interviews, calls, lectures, meetings, and other audio recordings.

Alibaba is paying more attention to the conversation itself. Its models can respond to interruptions, access external functions, and engage in voice dialogue.

At this point, there is no clear winner. Alibaba's solution has proven stronger in voice conversation tests, but it significantly lags behind OpenAI in response time.

🔒 Full answers are available for free after registration
Загрузка...
Ask uncomfortable questions

Discuss controversial topics and theories.

Compare arguments and opinions.

Compare arguments and opinions.

Top comments rise due to votes.

Top comments rise due to votes.

Consultant Online