Live data from Hacker News

Audiobox: Meta's new foundation research model for audio generation

ai.meta.com

71–80 of 90 posts

Re: Audiobox: Meta's new foundation research model for audio generation

#71

Earlier quoted context omitted.

I’m an amateur music producer and vocals are by far the toughest part of making music. I have to find a singer, convince them to work with me (I am an amateur and not particularly good tbh), and book studio space because its very tough to get a clean recording at home. I’m hoping that like digital instruments, I’ll be able to splice in digital voices instead of finding singers.

This is already somewhat available (check out Dreamtomics Synthesizer V and the voices like Solaris etc)

I have been using Synthesizer V for a lot of my music and the quality is very high, even with little manual tuning. There are a lot of voices to choose from now, and they added a cross-synthesis feature which lets you use the voices originally intended for Japanese and Chinese in English, and they don’t have a strong accent (I do find that I have to mess with the lyrics input a little). Also SAROS is coming out which has experimental Spanish support.

Definitely recommend looking up some Synthesizer V covers to see how realistic singing voice synthesis has become in recent years. It’s also free to try lower quality versions of the voices.

Re: Audiobox: Meta's new foundation research model for audio generation

#72
Regarding their "responsible" model, Meta's engineers aren't stupid. They know that:

1. No TTS audio output is tamper-proof. Their "safeguards" will be busted, and quickly. Whether via a small adversarial NN, some basic DSP, or just...holding a cheap recorder near your speakers, maintaining audio file provenance has no chance.

2. Impersonations have vexed humanity since the invention of vocal cords. Insofar as it's soluble, it's been solved -- authenticity is determined by a fluid mixture of context, trustworthiness, and the authority of involved parties & institutions. Always has been. Always will be. If I could drill one idea deep into every tech evangelist's head, it'd be: The solution to every problem isn't automatically "more technology." But hammers see only nails, so the vicious cycle continues, and society deals with the consequences (e.g. cryptobros decentralizing money...by slowly reinventing banks, but with more fraud).

3. This secret audio ID "feature" is probably harmful. It adds needless complexity. At best it exacerbates a false sense of safety because impersonation is trivial. Bad guys can emulate it on authentic recordings to discredit them as "fake." Nobody who'd actually benefit from such safeguards will respect them. News says this audio that affirms my confirmation bias is fake? Nah, the news is fake.

Meta knows all of this. Optimistically, I hope it's just lip service to concern fetishists; plausible deniability for the knife manufacturer when a bad guy uses one. Pessimistically, it might be pretext for an about-face on their OSS commitment. "Oops, researchers trivially broke our safeguards. Shucks. That's scary. Guess we'll build a moat instead of an OSS community. Think of the children or terminator or whatever works these days"

I suppose we'll see.

Re: Audiobox: Meta's new foundation research model for audio generation

#73
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

I'm old and I sympathize with your last sentence, I _imagine_ certain reinforced paths through repeated experience are myelin covered so much as to be almost static. But if I think of generations growing up, their brains aren't reinforced/seasoned enough to spot these "Wait... What?" moments in online content that is almost sensical but not quite. In a world with ever increasing AI hallucinated content, when children absorb fake content and build strong paths in their brains, at some point the words/meaning could be so distorted that you cannot understand anymore the person in front of you? I think of the political landscape where we have similar problems, social media algorithms curate content to keep you entertained within subjects A and B, and you have a community with shared values and you are seasoned and share its language domain. How would the landscape look like when the net can be bombarded by content that appears to be true, valid and useful but it really is not, for young generations? Also if somebody can paint the right myoline picture in my head (not plausible text generation but science) I would really appreciate it.

Re: Audiobox: Meta's new foundation research model for audio generation

#74
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

> That's a dream come true for every gamer on the face of this earth, I believe.

There's definitely some gamers who would like to never talk to a character in a video game, even if it means hours of bumbling around for the blue key that some NPC just told me about while I ignored 100% of the text in the game.

Re: Audiobox: Meta's new foundation research model for audio generation

#75
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

So on one hand there is the prospect of "complete narrative freedom for players" and on the other hand there is a fake information dystopia right ahead of us. Well, I would suspect, that the more dystopia we are going to have, the more people will want to flee that reality for a while into ever getting better role playing games. Is that, what people once called progress?

"a world where the only thing you can trust is people sitting right in front of you"

Oh and I think real people can be a source of missinformation as well. Also holograms(or brain implants) might have a breakthrough soon, as well. So all in all I think we are living in interesting times.

Re: Audiobox: Meta's new foundation research model for audio generation

#76

This a fantastic new development in the AI Audio space! However, it's quite disappointing that the model is closed sourced. Nonetheless, Alibaba's equivalent was released earlier in Nov and it's open-sourced! https://github.com/QwenLM/Qwen-Audio Does anyone have suggestions for how to integrate this into your tech stack via an internal API? Interested to hear the varying thoughts on this. From what I softly understan…

> Alibaba's equivalent was released earlier in Nov and it's open-sourced! https://github.com/QwenLM/Qwen-Audio

Openly distributed perhaps, but definitely not open source. The license appears as closed as meta (research-only with some leeway for other uses.)

Do you have any truely open-source general audio generation models yet?

I know about StyleTTS2, which is open source (MIT) and uncensored, but that model focuses on speech generation only. Having an non proprietary model like audiobox or Qwen-Audio would be really nice.

Re: Audiobox: Meta's new foundation research model for audio generation

#77

This a fantastic new development in the AI Audio space! However, it's quite disappointing that the model is closed sourced. Nonetheless, Alibaba's equivalent was released earlier in Nov and it's open-sourced! https://github.com/QwenLM/Qwen-Audio Does anyone have suggestions for how to integrate this into your tech stack via an internal API? Interested to hear the varying thoughts on this. From what I softly understan…

Normally I want basically everything to be open source. But as soon as I saw that audio restyling demo, I began to feel concerned they may be open sourcing this. The model can take a sample speakers voice, new text to speak, and also a description of a new location (like a cathedral with many echoes, or other background noises) and produce a convincing new audio sample.

This technology will present serious challenges for the verification of covertly recorded audio. It will of course ultimately become widespread but I’m not inherently bothered by the idea of slowing down its release. Giving researchers extra time to examine possible detection techniques seems helpful to me.

Re: Audiobox: Meta's new foundation research model for audio generation

#78

How long before someone manages to clone him/herself online and apply for relatively simple gig work. And duplicates than 1000 times. Making millions with simple work. Its almost possible i think. Clone your voice with this, clone your looks by wiring comfyui like sd nodes to your webcam. Everything instructed/orchestrated by some AI agent controlled by chatgpt. Some wiring logic is what you need to make. The only th…

By the time a machine learning model can replace 1000 workers they’ll just stop hiring real workers. What remains will be tasks which can’t be automated.

Re: Audiobox: Meta's new foundation research model for audio generation

#79
post #40

Earlier quoted context omitted.

This is why I'm high on the metaverse long term. In ten years, there will be a $500 (or whatever the 2033 inflation adjusted value is) VR headset that blows the Apple Vision Pro out of the water in terms of optics, will run a highly optimized version of the lastest revision of Llama locally (and it will be much better than anything we currently have today), come with wifi 8 (so it will have multigigabit per second re…

In 2040 we will have lenses with 32k displays and gpus with 10 trillion transistors and 1 petabyte of memory. Its hard to predict what you can do with that. The real world would be empty by then.

The real world is much richer than a bunch of stuff just being projected in to your eyes. You can’t climb a tree in VR.

Re: Audiobox: Meta's new foundation research model for audio generation

#80
post #73
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

I'm old and I sympathize with your last sentence, I _imagine_ certain reinforced paths through repeated experience are myelin covered so much as to be almost static. But if I think of generations growing up, their brains aren't reinforced/seasoned enough to spot these "Wait... What?" moments in online content that is almost sensical but not quite. In a world with ever increasing AI hallucinated content, when children…

I think that an answer to finding a 'true' reality on the net comes back to the centuries old dicipline of philosophy. The world has been full of 'garbage-thought' for ever. I might suggest that there is less 'garbage-thought' now for those that wish it than there ever has been.

The AI-Voice revolution leads us instead to the problem of authority, and of imitation of authority. Too many of us take it as given that an authority has the truth.

The AI-Voice / AI-Content revolution allows low quality actors to imitate relevant authorities.

So, to address your question, we need to study philosophy in schools, so that kids learn to think for themselves.

Post reply on HN