Live data from Hacker News

Viewing profile — abdik

abdik

HN member
Joined
Mon, Feb 06, 2017, 11:00 AM UTC
HN karma
67
Public activity
28 items

About abdik

No profile information was provided.

Recent public activity

  1. comment
    Comment #49350874

    not really counter-positioned, they are one of the providers we route to. elevenlabs sells models - voices, stt, and their agent platform on top of them. we don't sell any models: …

  2. comment
    Comment #49341524

    i don't have a thesis on voice as a form factor for search. the demand we serve already exists: businesses answer phones. receptionists, outbound campaigns, clinic front desks and …

  3. comment
    Comment #49341522

    thanks! yes, some of them do, cuz speed is a per-provider capability, not universal. And, you can see it in the gateway code (minimax, hume, xai tts adapters all handle a speed par…

  4. comment
    Comment #49341317

    People who built or building voice agents immediately get this. openrouter is the openrouter for audio models. the conflation is "audio models" vs "voice ai", and the mental model …

  5. comment
    Comment #49337991

    yes, it does support, when you are creating an api key, you can point out narration or transcription use case, then you will be able to see. let me know how it goes or what use cas…

  6. comment
    Comment #49336840

    looks cool, checking it out!

  7. comment
    Comment #49336388

    exactly, we see the same thing, around 95% cases are still cascaded, even tho STS has been improving a lot

  8. comment
    Comment #49336376

    Synthetic-data fine-tuning is the other credible answer to domain vocabulary. Curious whether you re-benchmark the fine-tune when new base models ship?

  9. comment
    Comment #49336354

    What purrcat259 said, and I think that the page should explain it, we will add a tooltip. for some languages CER is more relevant than WER. Thai and Mandarin have no word boundarie…

  10. comment
    Comment #49336281

    agree with omneity here. Whisper's initial-prompt trick is exactly that, and several hosted vendors have equivalents (custom vocabulary / keyword prompting). Domain vocabulary is w…

  11. comment
    Comment #49336259

    Actually, we measured exactly this recently. The strongest open-weights speech-to-speech model we have run is NVIDIA's NemotronLabs VoiceChat 11B - no provider serves it, so we hos…

  12. comment
    Comment #49336226

    cool app, and agreed that on-device keeps eating the single-user cases, we are seeing dictation and translation are exactly where local models shine. We benchmark the open models o…

  13. comment
    Comment #49336215

    Turn-taking specifically does not need a listening panel, but it is measured mechanically. 200+ real human clips, and we score end-vs-wait decisions: did the model decide the calle…

  14. comment
    Comment #49336185

    Yes, on the hosted side (agents platform): full sessions come with VAD and turn-taking handled - we set them up and tune them for your use case, so that is the closest thing to con…

  15. comment
    Comment #49335348

    there is a good progress on on-devise models, but not ready for production yet to fit in devices. But as soon as there is are some good results, we are going to benchmark them and …

  16. comment
    Comment #49334137

    thanks! actually, we have the filipino already, can you check out and share your feedback?

  17. comment
    Comment #49333994

    Fair pushback. On end to end: we measure those too, same methodology: https://benchmarks.speko.ai/s2s . If the single models win, we route to them the same way, so we do not care w…

  18. comment
    Comment #49333622

    The main difference from gateway is we help with picking the right voice stack, which seems to be a big problem for users: we benchmark the models continuously and route based on t…

  19. comment
    Comment #49333425

    Thank you!

  20. comment
    Comment #49333388

    Yes, i added it. that's the right link.

  21. story
    Launch HN: Speko (YC S26) – OpenRouter for Voice AI

    Hi HN! I'm Bek, founder of Speko, a platform that finds an optimal combination of speech-to-text, LLM, and text-to-speech models, given your constraints, among all our public bench…

  22. story
    Show HN: YC Interview Simulator (Voice of Garry Tan)

    Got into YC interview recently and built this to practice, it's been helpful. Sharing, so it can be used by others well. Good luck everyone!

  23. comment
    Comment #29290560

    The logo is similar to ours https://www.lovo.ai/

  24. story
  25. comment
    Comment #29127461

    Check out https://nomadlist.com/