OpenAI & Alibaba's New Voice AI Models for Developers

4d ago·0:00 listen·Source: dev.ua

Summary

OpenAI and Alibaba have simultaneously released new voice AI models for developers. These updates focus on improving voice processing and generation. OpenAI expanded its API with two new transcription tools: GPT-Live-Transcribe and GPT-Transcribe. GPT-Live-Transcribe is for real-time tasks, while GPT-Transcribe handles pre-recorded audio. Both can consider context like industry terms and expected languages. This context support increased semantic accuracy for GPT-Live-Transcribe from 38.5% to 44.6%, and for GPT-Transcribe from 41.6% to 45.2%. Their error rates also dropped significantly in tests. Meanwhile, Alibaba introduced the Qwen-Audio-3.0-Realtime Plus and Flash models, designed for direct speech-to-speech interaction. These support full-duplex mode, letting users interrupt the AI, and allow for external functions and cloned voices. The Qwen-Audio-3.0-Realtime Plus model scored 84.1% on the Speech-to-Speech Index, surpassing OpenAI's GPT-Realtime-2.1 High at 79.1%. However, Alibaba's model has a latency of about 4 seconds for the first audio signal, compared to OpenAI's 1.14 seconds. These advancements mean AI voice translation is moving closer to natural communication, preserving emotions and intonations.

Read the full article on dev.ua

This is an AI-generated audio summary. Always check the original source for complete reporting.

Share
Keep Listening