Live data from Hacker News

Viewing profile — jpcl

jpcl

HN member
Joined
Fri, Oct 26, 2018, 9:16 PM UTC
HN karma
82
Public activity
32 items

About jpcl

No profile information was provided.

Recent public activity

  1. comment
    Comment #47351202

    What this means is that it does not support things like acting instructions or creating a voice from a text description. If you prompt it with a matching text+voice sample it will …

  2. comment
    Comment #39130421

    Hi, I used the [WhisperSpeech]( https://github.com/collabora/WhisperSpeech ) model for the TTS part after I did some serious torch.compile optimizations to bring the latency down. …

  3. comment
    Comment #39054450

    Not really, the Polish effort is run by a non-profit and hired professional voice actors.

  4. comment
    Comment #39053721

    Yeah, thanks. I'd love to try to clarify this (I'll put this into our documentation as well ASAP) for anyone that may be reading this in the future: Our model is not a derivative o…

  5. comment
    Comment #39045480

    Did exactly that, thanks for spotting that. :) https://github.com/collabora/WhisperSpeech/commit/398b889060...

  6. comment
    Comment #39045343

    Good point, thanks. And I was thinking it will show that the model can really synthesize very varied samples... ;)

  7. comment
    Comment #39045313

    That's an interesting thought. The semantic tokens we get from Whisper serve a similar purpose – you can convert existing speech to different voices, I did not try with accents yet…

  8. comment
    Comment #39044858

    That's true but you make it sound like it's totally obvious where the line of fair use should be drawn for AI training. Until courts or lawmakers make it clearer I personally belie…

  9. comment
    Comment #39044790

    We looked at this at one point and it seems whisper.cpp/llama.cpp have all the bits needed to make it work. I'd love to help if someone wanted to give it a shot.

  10. comment
    Comment #39044451

    https://wolnelektury.pl/katalog/audiobooki/ is the Polish audiobook collection. The English audiobooks are public domain recordings from LibriVox (via the LibriLight dataset).

  11. comment
    Comment #39044422

    Thanks, I'll check it out. I don't know any Chinese so I'll probably reach out to you for some help :)

  12. comment
    Comment #39044396

    Yeah, we'd love to help you when you decide to give it a try so feel free to reach out. We also have quite a few people working on VR at Collabora.

  13. comment
    Comment #39044349

    We are working hard to uphold all the licensing rules but nobody can absolve you from all legal risks. There may be a court ruling/new law that any training needs a special permiss…

  14. comment
    Comment #39042136

    For Polish I have around 700hr. I suspect that we will need less hours if we add more languages since they do overlap to some extent. Fixed transcripts would be nice although we ne…

  15. comment
    Comment #39042080

    Yes, you download the weights once from Huggingface and you can do whatever you want with it. :) We have no cloud APIs or usage tracking of any kind.

  16. comment
    Comment #39042047

    Both models are using around 3GB right now (converted into FP16 for speed). But I checked that the (slower) FP32 version uses 2.3GB so we are probably doing something suboptimal he…

  17. comment
    Comment #39041288

    There is a plot of language performance on their repo: https://github.com/openai/whisper I am not aware of a multi-lingual leaderboard for speech recognition models.

  18. comment
    Comment #39041264

    Training the T2S model from scratch takes around 8h on 96 A100 GPUs. Training the `tiny` S2A model is around 3x faster (training HQ `small` variant is comparable to T2S). I think y…

  19. comment
    Comment #39040575

    Yeah, Whisper is not clear-cut but since it is not a generative model I think their data usage is a lot more likely to be considered fair-use. And the part of that which we use for…

  20. comment
    Comment #39040562

    Hi, WhisperSpeech dev here. Thanks for all the nice comments, I was working really hard on this model for quite a few months now but there are still a lot of ways we can make it be…

  21. comment
    Comment #39040536

    Hi, thanks a lot for the tip, I'll update the README samples ASAP. :) I was busy working on inference performance in the last few weeks and totally did not expect to land on Hacker…

  22. comment
    Comment #39040518

    Both Polish and English samples are actually synthesized with a voice trained on the WolneLektury audiobooks. They are the highest quality open source (CC BY-SA) audiobooks I could…

  23. comment
    Comment #39040490

    Yeah, the Mimic is a lot less resource intensive. We are working to improve WhisperSpeech in this regard but it's probably always going to require more compute (but in return you'l…

  24. comment
    Comment #39040472

    Yup, we are using Whisper to transcribe automatically so we can train the model on just speech recordings, without human transcripts. This works for any language that is well suppo…

  25. comment
    Comment #39040460

    Thanks a lot. :) We are constantly working on these models and we push new versions every two months or so. It should get even better soon. :)