OpenAI launches advanced audio models for voice agents

Published 2025-03-20, 02:50 p/m
© Reuters

Investing.com -- OpenAI has unveiled a new suite of advanced audio models designed to enhance voice agents with improved speech-to-text and text-to-speech capabilities. These models are now available to developers worldwide, building upon OpenAI’s previous agent technologies like Operator and Deep Research.

The new speech-to-text models, gpt-4o-transcribe and gpt-4o-mini-transcribe, set new benchmarks for accuracy even in challenging conditions with noise, accents, and varying speech speeds. They show improved Word Error Rate performance compared to existing Whisper models, making them ideal for applications like call centers and meeting transcription.

For text-to-speech, the new gpt-4o-mini-tts model offers unprecedented "steerability," allowing developers to specify how content should be spoken. Developers can now create voice agents that speak like sympathetic customer service representatives or expressive storytellers, though currently limited to preset artificial voices.

These audio models were extensively pretrained on specialized audio datasets using the GPT-4o architectures to optimize performance. OpenAI has also enhanced its distillation techniques to transfer knowledge from larger models to smaller, more efficient ones.

The models are accessible through APIs with simplified integration options for developers already working with text-based models. OpenAI plans to continue improving these technologies while exploring custom voice options and expanding into other modalities like video for more personalized multimodal experiences.

Latest comments

Risk Disclosure: Trading in financial instruments and/or cryptocurrencies involves high risks including the risk of losing some, or all, of your investment amount, and may not be suitable for all investors. Prices of cryptocurrencies are extremely volatile and may be affected by external factors such as financial, regulatory or political events. Trading on margin increases the financial risks.
Before deciding to trade in financial instrument or cryptocurrencies you should be fully informed of the risks and costs associated with trading the financial markets, carefully consider your investment objectives, level of experience, and risk appetite, and seek professional advice where needed.
Fusion Media would like to remind you that the data contained in this website is not necessarily real-time nor accurate. The data and prices on the website are not necessarily provided by any market or exchange, but may be provided by market makers, and so prices may not be accurate and may differ from the actual price at any given market, meaning prices are indicative and not appropriate for trading purposes. Fusion Media and any provider of the data contained in this website will not accept liability for any loss or damage as a result of your trading, or your reliance on the information contained within this website.
It is prohibited to use, store, reproduce, display, modify, transmit or distribute the data contained in this website without the explicit prior written permission of Fusion Media and/or the data provider. All intellectual property rights are reserved by the providers and/or the exchange providing the data contained in this website.
Fusion Media may be compensated by the advertisers that appear on the website, based on your interaction with the advertisements or advertisers.
© 2007-2025 - Fusion Media Limited. All Rights Reserved.