Launch HN: Speko (YC S26) – OpenRouter for Voice AI
11–20 of 73 posts
Re: Launch HN: Speko (YC S26) – OpenRouter for Voice AI
#12> Typical production voice agent is an ensemble of three models: STT, an LLM, and TTS. To use a claudism , I would like to push back on this. The industry is very much moving towards one-model-does-all end to end trained similar to LLMs and VLMs. Mostly for latency reasons and partially because the results for the end to end trained models are just so much better than those using three pieces architectures. I think m…
On "promptfoo of voice models": that is closer to how it started. At my last company we ran these evals manually, we would even hire native-speaking raters, benchmark, switch if it wins. The evals are the value, agreed. The routing is what makes them actionable: teams told us swapping always looked like an R&D project, so scores alone did not change what ran in production.
On prompt-based voice gen and reference-audio cloning: agreed, that is what we see too. It makes continuous measurement more important: the same style prompt behaves differently per language and per content type, so we rank the voices themselves, tagged by use case: https://benchmarks.speko.ai/tts-voices
Re: Launch HN: Speko (YC S26) – OpenRouter for Voice AI
#13Re: Launch HN: Speko (YC S26) – OpenRouter for Voice AI
#14> Typical production voice agent is an ensemble of three models: STT, an LLM, and TTS. To use a claudism , I would like to push back on this. The industry is very much moving towards one-model-does-all end to end trained similar to LLMs and VLMs. Mostly for latency reasons and partially because the results for the end to end trained models are just so much better than those using three pieces architectures. I think m…
Re: Launch HN: Speko (YC S26) – OpenRouter for Voice AI
#15Looks awesome, can't wait to try it for some Filipino workflows when it's available!
Re: Launch HN: Speko (YC S26) – OpenRouter for Voice AI
#16> Typical production voice agent is an ensemble of three models: STT, an LLM, and TTS. To use a claudism , I would like to push back on this. The industry is very much moving towards one-model-does-all end to end trained similar to LLMs and VLMs. Mostly for latency reasons and partially because the results for the end to end trained models are just so much better than those using three pieces architectures. I think m…
Re: Launch HN: Speko (YC S26) – OpenRouter for Voice AI
#17Re: Launch HN: Speko (YC S26) – OpenRouter for Voice AI
#18Re: Launch HN: Speko (YC S26) – OpenRouter for Voice AI
#19Seems to be useless, the state of art for all categories is local on-device, voice model vendors are just rent seekers for those who know no better.
Re: Launch HN: Speko (YC S26) – OpenRouter for Voice AI
#20Seems to be useless, the state of art for all categories is local on-device, voice model vendors are just rent seekers for those who know no better.
Just canceled my WisprFlow subscription a few days ago to switch to a open source, free, local alternative. In my case handy.computer with the Cohere model.