Live data from Hacker News

Viewing profile — jeffharris

jeffharris

HN member
Joined
Tue, Jul 07, 2015, 4:42 PM UTC
HN karma
99
Public activity
22 items

About jeffharris

No profile information was provided.

Recent public activity

  1. comment
    Comment #43430358

    oh doh. thanks ... we just pushed a fix for the crash. Unfortunately our currently implementation needs service works for streaming audio, so the "fix" was to disable the feature i…

  2. comment
    Comment #43430310

    really depends on the voice. but in general, no we want to sound as realistic as possible and I expect future voices to keep improving on this front

  3. comment
    Comment #43430288

    this is a good solve. we don't support word time stamps natively yet, but are working on teaching GPT-4o that skill

  4. comment
    Comment #43430260

    S2S is where we're investing the most effort on audio ... sorry it's been slow but we are working hard on it Top priorities at the moment 1) Better function calling performance 2) …

  5. comment
    Comment #43430241

    we're working hard on it at the moment and hope we'll have a snapshot ready in the next month or so we've debugged the cutoff issues and have fixes for them internally but we need …

  6. comment
    Comment #43430209

    this has been coming up often recently. nothing to announce yet, but when enough developers ask for it, we'll build it into the model's training diarization is also a feature we pl…

  7. comment
    Comment #43430193

    we'll keep expanding these GPT-4o based models with more controls. Is the main feature missing we're missing custom voices?

  8. comment
    Comment #43427747

    not open source at this time. unfortunately they're much to large to run on normal consumer hardware

  9. comment
    Comment #43427738

    Should do! here's an example https://www.openai.fm/#4a5a82db-faea-4f80-813c-3131902c2458

  10. comment
    Comment #43427700

    thanks for flagging ... number fidelity (especially on languages that are unfortunately less represented in training data) is still something we're working to improve

  11. comment
    Comment #43427579

    nothing to share on open source yet, it's something we'll keep exploring. Especially as the models get smaller so more able to run on regular devices

  12. comment
    Comment #43427557

    try the ballad or fable voices

  13. comment
    Comment #43427550

    We've been using the FLUERS eval and you can see comparisons to other models on the market in the post https://openai.com/index/introducing-our-next-generation-aud... Curious if th…

  14. comment
    Comment #43427525

    It's a slightly better model for TTS. With extra training focusing on reading the script exactly as written. e.g. the audio-preview model when given instruction to speak "What is t…

  15. comment
    Comment #43427503

    1/ we've been working a lot on accents, so expect improvements with these models... though we're not done. Would be curious how you find them. And try giving specific detailed inst…

  16. comment
    Comment #43426980

    Yes, from our terms: "Don’t build tools that may be inappropriate for minors, including: Sexually explicit or suggestive content. This does not include content created for scientif…

  17. comment
    Comment #43426970

    We're thinking about diarization (adding time awareness to GPT models) but no firm plans to share just yet

  18. comment
    Comment #43426636

    so good! https://www.openai.fm/#28540f27-5b51-445a-b1d6-1c89711a2c4f

  19. comment
    Comment #43426603

    some of the older voices are definitely less steerable, more robotic we put little stars in the bottom right corner for the newer voices, which should sound better

  20. comment
    Comment #43426578

    Hey, I'm Jeff and I was PM for these models at OpenAI. Today we launched three new state-of-the-art audio models. Two speech-to-text models—outperforming Whisper. A new TTS model—y…

  21. comment
    Comment #41640043

    apologies: it's taken us a minute to switch the default `gpt-4o` pointer to the newest snapshot we're planning on doing that default change next week (October 2nd). And you can get…

  22. comment
    Comment #41174929

    yep image input on the new model is also 50% cheaper and apologies for the outdated pricing calculator ... we'll be updating it later today