Hey, I'm Jeff and I was PM for these models at OpenAI. Today we launched three new state-of-the-art audio models. Two speech-to-text models—outperforming Whisper. A new TTS model—you can instruct it how to speak (try it on openai.fm!). And our Agents SDK now supports audio, making it easy to turn text agents into voice agents. We think you'll really like these models. Let me know if you have any questions here!
> Two speech-to-text models—outperforming Whisper On what metric? Also Whisper is no longer state of the art in accuracy, how does it compare to the others in this benchmark? https://artificialanalysis.ai/speech-to-text
Curious if there's a benchmark you trust most?