Live data from Hacker News

Audiobox: Meta's new foundation research model for audio generation

ai.meta.com

41–50 of 90 posts

Re: Audiobox: Meta's new foundation research model for audio generation

#41

I think the release of closed source models right now is a net negative and worth opposing. Right now we’re building a future where the very wealthy and powerful will control access to AI on ethical grounds, while they have uncensored access to the latest and most powerful models. Innovation, high frequency trading, medical breakthroughs, creative output - all of these and more will be enhanced by AI, and you’ll be e…

I mean sure.. but imagine if this were open sourced as is. This is new tech that has barely had time to mature. The possibilities for abuse are endless. I for one am happy that this model isn't being open sourced. This is an excellent way for people to generate all kinds of disturbing and fake audio clips.

Re: Audiobox: Meta's new foundation research model for audio generation

#42
post #33

Earlier quoted context omitted.

And the millions of other RVC websites. Musicfy, Uberduck, Coversai, Kitsai, FakeYou, Voicemyai, Voicify, Bangerapp, Tryreplay, Weightsgg ... RVC is so easy anyone can spin up a website for it. No moat. Over a hundred thousand trained weights files in the open, so it's easy to bootstrap.

For anybody else who also hadn't come across the term RVC before: "The RVC model is a Retrieval-based Voice Conversion system using AI for high-quality voice cloning. It utilizes artificial intelligence to modify or clone voices in real-time." Source: https://speechify.com/blog/rvc-vocal-models/

Wow. I'm going to give each family member a password and make them prove that they're real when they call me :)

Re: Audiobox: Meta's new foundation research model for audio generation

#43

I think the release of closed source models right now is a net negative and worth opposing. Right now we’re building a future where the very wealthy and powerful will control access to AI on ethical grounds, while they have uncensored access to the latest and most powerful models. Innovation, high frequency trading, medical breakthroughs, creative output - all of these and more will be enhanced by AI, and you’ll be e…

[deleted]

Re: Audiobox: Meta's new foundation research model for audio generation

#44
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

> react to your words and actions in a believable, fully immersive manner “I’m sorry, but as an ethically trained AI I cannot engage in this sword fight. Violence is never the answer.” Yeah. It’s gonna be very immersive :P

[deleted]

Re: Audiobox: Meta's new foundation research model for audio generation

#46
post #5

Earlier quoted context omitted.

Sounds cool - got any specific ones to share?

https://youtu.be/A7tp4eg0ax8

The new avatar (blue skin not arrow) game looks like this demo with some characters tossed in to control.

Re: Audiobox: Meta's new foundation research model for audio generation

#47
post #33
post #31

Earlier quoted context omitted.

This already exists, ex. Audimee: https://audimee.com/

And the millions of other RVC websites. Musicfy, Uberduck, Coversai, Kitsai, FakeYou, Voicemyai, Voicify, Bangerapp, Tryreplay, Weightsgg ... RVC is so easy anyone can spin up a website for it. No moat. Over a hundred thousand trained weights files in the open, so it's easy to bootstrap.

Great, so the competition is going to yield cheap services with a nice UX.

Re: Audiobox: Meta's new foundation research model for audio generation

#48
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

The amount of context needed would require quite a bit of novel R&D that doesn't exist on the horizon yet. I think it's more realistic that it'll be a mixture of real & fake in the interim (e.g. the LLM will record important game state changes / information & then use that as context but it'll still forget a bunch of things you'd expect a human to).

Re: Audiobox: Meta's new foundation research model for audio generation

#49
post #41

I think the release of closed source models right now is a net negative and worth opposing. Right now we’re building a future where the very wealthy and powerful will control access to AI on ethical grounds, while they have uncensored access to the latest and most powerful models. Innovation, high frequency trading, medical breakthroughs, creative output - all of these and more will be enhanced by AI, and you’ll be e…

I mean sure.. but imagine if this were open sourced as is. This is new tech that has barely had time to mature. The possibilities for abuse are endless. I for one am happy that this model isn't being open sourced. This is an excellent way for people to generate all kinds of disturbing and fake audio clips.

Same logic could be applied to Linux by Microsoft in the 90s. In fact, the “It’s for your safety” has been applied to some of the worst things humans have perpetrated including apartheid (which I lived through) and the holocaust. And it’s always those that claim to keep us safe doing the worst. And it continues with perceived dangers providing pretext and moral authority to do bad things.

Re: Audiobox: Meta's new foundation research model for audio generation

#50
post #40
post #2

VR is gonna get wild in like 5 years if they keep this up

This is why I'm high on the metaverse long term. In ten years, there will be a $500 (or whatever the 2033 inflation adjusted value is) VR headset that blows the Apple Vision Pro out of the water in terms of optics, will run a highly optimized version of the lastest revision of Llama locally (and it will be much better than anything we currently have today), come with wifi 8 (so it will have multigigabit per second re…

And yet I'll still be waiting five weeks to just download the world because I'm stuck on 2Mb/s ADSL.
Post reply on HN