We benchmarked OpenAI's newest gpt-4o-transcribe model against JigsawStack STT across real-world scenarios, including performance, accuracy, feature set, and multilingual capabilities. Here's what we found:
Performance:
- JigsawStack processes audio ~2.4x faster than OpenAI's model across all audio lengths
- This speed advantage remains consistent with both short clips and longer audio files
File Support:
- OpenAI limits files to 25MB and 25 minutes of audio per request
- JigsawStack handles files up to 100MB and 4 hours of audio per request
Advanced Features:
- Speaker Recognition: JigsawStack accurately identifies up to 50 different speakers with timestamps
- Detailed Timestamps: Provides sentence-level timing data by default
- Translation: Built-in support for translating audio into 100+ languages
Real-world Accuracy:
- Achieved 100% accuracy on noisy audio samples where OpenAI scored 94%
- Properly identified repeated phrases and sentence boundaries that OpenAI missed
Our team has focused on building a specialized solution that outperforms in nearly every metric that matters for real-world applications.
Check out the full breakdown here: http://jigsawstack.com/blog/openai-audio-stt-vs-jigsawstack-...
Let us know what you think! We'd love feedback from the HN community - particularly from those working with audio transcription at scale.