you're missing a lot
TTS: 11labs, PlayHT, Cartesia, iFLYTEK, AWS Polly, Deepgram Aura
STT: Deepgram (multiple models, including Whisper), Gladia Whisper, Soniox
just off the top of my head (it's my dayjob!)
51–60 of 82 posts
you're missing a lot
TTS: 11labs, PlayHT, Cartesia, iFLYTEK, AWS Polly, Deepgram Aura
STT: Deepgram (multiple models, including Whisper), Gladia Whisper, Soniox
just off the top of my head (it's my dayjob!)
Great! I wish there was a "bang to buck" value. Some way to know the cheapest model I could use for creating structured data from unstructured text, reliably. Using gpt4o-mini which is cheap but wouldn't know if anything cheaper could do the job too.
I always plug openrouter.ai for making cross-model comparisons. It's my general goto for random stuff. (I am not affiliated, just a user)
OP, were you inspired by this LLM comparison tool? https://whatllm.vercel.app The tables are very similar - though you've added a custom calculator which is a nice touch. Also for the Versus Comparison, it might be nice to have a checkbox that when clicked highlights the superlative fields of each LLM at a glance.
This page has up to date information of all models and providers: https://artificialanalysis.ai/leaderboards/providers We also on other pages cover Speech to Text, Text to Speech, Text to Image, Text to Video.
Note I'm one of the creators of Artificial Analysis.
I like the idea of more comparisons of models. Are there plans to add independent analyses of these models or is it only an aggregation of input limits? How do you see this differing from or adding to other analyses such as: https://artificialanalysis.ai https://huggingface.co/spaces/TTS-AGI/TTS-Arena https://huggingface.co/spaces/hf-audio/open_asr_leaderboard https://huggingface.co/spaces/TIGER-Lab/GenAI-Arena Great…
the gradio ui looks ugly imo, that's why I used shadcn and next.js to make the website look good. I'll try to make it as user-friendly as possible. Most of the websites are ugly + too technical.
I like your work visually on first glance, god knows you're right about gradio, even if its irrelevant.
But peddling extremely limited, out of date, versions of other people's data, trumps that, especially with this tagline. "A website to compare every AI model: LLMs, TTSs, STTs"
It is a handful of LLMs, then one TTS model, then one STT model, both with 0 data. And it's worth pointing out, since this endeavor is motivated by design trumping all: all the columns are for LLM data.
In my own experiments with the chat models they seem to lose the plot after about 10 replies unless constantly "refreshed", which is a tiny fraction of the supposed 128000 token input length that 4o has. Does Gemini actually do something dramatically differently, or is their 3 million token context window pure marketing nonsense?
As far as I know, there's a volcano engine in China that has impressive text-to-speech capabilities. Many local companies are using this model.