I think the release of closed source models right now is a net negative and worth opposing. Right now we’re building a future where the very wealthy and powerful will control access to AI on ethical grounds, while they have uncensored access to the latest and most powerful models. Innovation, high frequency trading, medical breakthroughs, creative output - all of these and more will be enhanced by AI, and you’ll be e…
Audiobox: Meta's new foundation research model for audio generation
51–60 of 90 posts
Re: Audiobox: Meta's new foundation research model for audio generation
#52This a fantastic new development in the AI Audio space! However, it's quite disappointing that the model is closed sourced. Nonetheless, Alibaba's equivalent was released earlier in Nov and it's open-sourced! https://github.com/QwenLM/Qwen-Audio Does anyone have suggestions for how to integrate this into your tech stack via an internal API? Interested to hear the varying thoughts on this. From what I softly understan…
and important, if you have more than 100m active users:) 4. Restrictions If you are commercially using the Materials, and your product or service has more than 100 million monthly active users, You shall request a license from Us
So, looks like it's absolutely fine to use, except for IT behemoth.
As for how to use, API, I think. Interesting applications are possible. Like interactive mobile robots. Assistants for people with disabilities, both software and wearable.
Interesting times... this will be called AI revolution probably. It's already not a joke, after several ups and downs.
Re: Audiobox: Meta's new foundation research model for audio generation
#53Earlier quoted context omitted.
For anybody else who also hadn't come across the term RVC before: "The RVC model is a Retrieval-based Voice Conversion system using AI for high-quality voice cloning. It utilizes artificial intelligence to modify or clone voices in real-time." Source: https://speechify.com/blog/rvc-vocal-models/
Wow. I'm going to give each family member a password and make them prove that they're real when they call me :)
Would make a nice vignette in a film about a dystopian future where video can be be generated cheaply and of sufficient quality.
Re: Audiobox: Meta's new foundation research model for audio generation
#54Earlier quoted context omitted.
This already exists, ex. Audimee: https://audimee.com/
And the millions of other RVC websites. Musicfy, Uberduck, Coversai, Kitsai, FakeYou, Voicemyai, Voicify, Bangerapp, Tryreplay, Weightsgg ... RVC is so easy anyone can spin up a website for it. No moat. Over a hundred thousand trained weights files in the open, so it's easy to bootstrap.
Re: Audiobox: Meta's new foundation research model for audio generation
#55Earlier quoted context omitted.
I mean sure.. but imagine if this were open sourced as is. This is new tech that has barely had time to mature. The possibilities for abuse are endless. I for one am happy that this model isn't being open sourced. This is an excellent way for people to generate all kinds of disturbing and fake audio clips.
Same logic could be applied to Linux by Microsoft in the 90s. In fact, the “It’s for your safety” has been applied to some of the worst things humans have perpetrated including apartheid (which I lived through) and the holocaust. And it’s always those that claim to keep us safe doing the worst. And it continues with perceived dangers providing pretext and moral authority to do bad things.
"I just hold on to all the money, 'cause bitches can't be trusted with it. We pool all the kissing money together, see? But if you wanna buy anything, you just talk to the bottom bitch, and then the bottom bitch talks to me.
Do you know what I am saying?"
Re: Audiobox: Meta's new foundation research model for audio generation
#56VR is gonna get wild in like 5 years if they keep this up
Re: Audiobox: Meta's new foundation research model for audio generation
#57Earlier quoted context omitted.
This is why I'm high on the metaverse long term. In ten years, there will be a $500 (or whatever the 2033 inflation adjusted value is) VR headset that blows the Apple Vision Pro out of the water in terms of optics, will run a highly optimized version of the lastest revision of Llama locally (and it will be much better than anything we currently have today), come with wifi 8 (so it will have multigigabit per second re…
And yet I'll still be waiting five weeks to just download the world because I'm stuck on 2Mb/s ADSL.
Re: Audiobox: Meta's new foundation research model for audio generation
#58Does Meta usually provide a web interface for them or do you have to download and run locally?
Re: Audiobox: Meta's new foundation research model for audio generation
#59Earlier quoted context omitted.
I’m an amateur music producer and vocals are by far the toughest part of making music. I have to find a singer, convince them to work with me (I am an amateur and not particularly good tbh), and book studio space because its very tough to get a clean recording at home. I’m hoping that like digital instruments, I’ll be able to splice in digital voices instead of finding singers.
This already exists, ex. Audimee: https://audimee.com/
Re: Audiobox: Meta's new foundation research model for audio generation
#60If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…
> react to your words and actions in a believable, fully immersive manner “I’m sorry, but as an ethically trained AI I cannot engage in this sword fight. Violence is never the answer.” Yeah. It’s gonna be very immersive :P
I could see Microsoft making the first move next generation since they're knee deep in it.