Live data from Hacker News

Audiobox: Meta's new foundation research model for audio generation

ai.meta.com

31–40 of 90 posts

Re: Audiobox: Meta's new foundation research model for audio generation

#31
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

I’m an amateur music producer and vocals are by far the toughest part of making music. I have to find a singer, convince them to work with me (I am an amateur and not particularly good tbh), and book studio space because its very tough to get a clean recording at home. I’m hoping that like digital instruments, I’ll be able to splice in digital voices instead of finding singers.

This already exists, ex. Audimee: https://audimee.com/

Re: Audiobox: Meta's new foundation research model for audio generation

#33
post #31

Earlier quoted context omitted.

I’m an amateur music producer and vocals are by far the toughest part of making music. I have to find a singer, convince them to work with me (I am an amateur and not particularly good tbh), and book studio space because its very tough to get a clean recording at home. I’m hoping that like digital instruments, I’ll be able to splice in digital voices instead of finding singers.

This already exists, ex. Audimee: https://audimee.com/

And the millions of other RVC websites. Musicfy, Uberduck, Coversai, Kitsai, FakeYou, Voicemyai, Voicify, Bangerapp, Tryreplay, Weightsgg ...

RVC is so easy anyone can spin up a website for it. No moat. Over a hundred thousand trained weights files in the open, so it's easy to bootstrap.

Re: Audiobox: Meta's new foundation research model for audio generation

#34
This a fantastic new development in the AI Audio space! However, it's quite disappointing that the model is closed sourced. Nonetheless, Alibaba's equivalent was released earlier in Nov and it's open-sourced! https://github.com/QwenLM/Qwen-Audio

Does anyone have suggestions for how to integrate this into your tech stack via an internal API? Interested to hear the varying thoughts on this. From what I softly understand is that the model weights have to be swapped or altered per se to be able to commercially reuse this. Correct me if I'm wrong.

Re: Audiobox: Meta's new foundation research model for audio generation

#35
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

I used to be an LLM until I took an arrow to the knee. Aside from the joke, I think the barrier would definitely reduce and in-game characters would be contextually far more aware but how do you enforce plot progression in such a truly open world? Can you control the boundary of LLM expression?

Re: Audiobox: Meta's new foundation research model for audio generation

#36
post #33
post #31

Earlier quoted context omitted.

This already exists, ex. Audimee: https://audimee.com/

And the millions of other RVC websites. Musicfy, Uberduck, Coversai, Kitsai, FakeYou, Voicemyai, Voicify, Bangerapp, Tryreplay, Weightsgg ... RVC is so easy anyone can spin up a website for it. No moat. Over a hundred thousand trained weights files in the open, so it's easy to bootstrap.

For anybody else who also hadn't come across the term RVC before:

"The RVC model is a Retrieval-based Voice Conversion system using AI for high-quality voice cloning. It utilizes artificial intelligence to modify or clone voices in real-time." Source: https://speechify.com/blog/rvc-vocal-models/

Re: Audiobox: Meta's new foundation research model for audio generation

#37
I think the release of closed source models right now is a net negative and worth opposing. Right now we’re building a future where the very wealthy and powerful will control access to AI on ethical grounds, while they have uncensored access to the latest and most powerful models. Innovation, high frequency trading, medical breakthroughs, creative output - all of these and more will be enhanced by AI, and you’ll be eating leftovers and paying a fortune for them, wondering why you can’t keep up - unless we enable a vibrant open source ecosystem, and force big tech to release models into that ecosystem.

Support open source models by celebrating their release and pressuring companies to release them, and oppose closed source AI or face a very bleak future for you and your descendants.

You may be having fun with “Open” AI’s API today, but you’re supporting and celebrating the collapse of society into megacap AI elites and a majority paying for metered access to old technology.

Re: Audiobox: Meta's new foundation research model for audio generation

#38
post #20

Earlier quoted context omitted.

Save your money and record at home under a blanket.

Layered towels are surprisingly capable too

Or if you want to feel a little fancy, a 3-sided folding project board with foam glued to it.

Re: Audiobox: Meta's new foundation research model for audio generation

#39
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

It would be fantastic to put a bunch of different LLMs in a game map with “senses” fulfilled by multimodal inputs and agency to carry out actions within the game’s universe. With a goal such as make the most money or rule the most kingdoms, it would be super interesting to see how it self organizes.

Re: Audiobox: Meta's new foundation research model for audio generation

#40
post #2

VR is gonna get wild in like 5 years if they keep this up

This is why I'm high on the metaverse long term. In ten years, there will be a $500 (or whatever the 2033 inflation adjusted value is) VR headset that blows the Apple Vision Pro out of the water in terms of optics, will run a highly optimized version of the lastest revision of Llama locally (and it will be much better than anything we currently have today), come with wifi 8 (so it will have multigigabit per second real word performance), capable of rendering graphics that look much more realistic than Unreal Engine 5 (with high frame rates due to AI upscaling and frame generation).

There will be people that will spend almost every waking hour with one of those things attached to their face if they can also make this device lightweight and comfortable

Post reply on HN