Live data from Hacker News

Audiobox: Meta's new foundation research model for audio generation

ai.meta.com

81–90 of 90 posts

Re: Audiobox: Meta's new foundation research model for audio generation

#81
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

So on one hand there is the prospect of "complete narrative freedom for players" and on the other hand there is a fake information dystopia right ahead of us. Well, I would suspect, that the more dystopia we are going to have, the more people will want to flee that reality for a while into ever getting better role playing games. Is that, what people once called progress? "a world where the only thing you can trust is…

> Oh and I think real people can be a source of missinformation as well.

Yeah, definitely. But at least you have a chance to detect their intentions; for some forged digital information, you don't even have that. "Interesting times" is one way to put it...

Re: Audiobox: Meta's new foundation research model for audio generation

#82

Earlier quoted context omitted.

In 2040 we will have lenses with 32k displays and gpus with 10 trillion transistors and 1 petabyte of memory. Its hard to predict what you can do with that. The real world would be empty by then.

The real world is much richer than a bunch of stuff just being projected in to your eyes. You can’t climb a tree in VR.

Maybe for some. But for those living in unfortunate circumstances (low socioeconomic status, small apartment, poor future prospects, etc) I bet a VR world of paradise and social connection is far more inviting.

Re: Audiobox: Meta's new foundation research model for audio generation

#83

How long before someone manages to clone him/herself online and apply for relatively simple gig work. And duplicates than 1000 times. Making millions with simple work. Its almost possible i think. Clone your voice with this, clone your looks by wiring comfyui like sd nodes to your webcam. Everything instructed/orchestrated by some AI agent controlled by chatgpt. Some wiring logic is what you need to make. The only th…

By the time a machine learning model can replace 1000 workers they’ll just stop hiring real workers. What remains will be tasks which can’t be automated.

Or isn't allowed to by law, e.g. security etc.

Re: Audiobox: Meta's new foundation research model for audio generation

#84
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

It'll affect linear entertainment far before it impacts games seriously (Especially the 'full immersive' games where you chat with AI agents) AI is still too expensive and performance intensive to run in games cost effectively, and truly powerful AI is probably another 10-100x cost increase. On the other hand, novels will be rapidly replaced by visual novels. The cost of having a novel fully illustrated and voiced wi…

Are you sure of this? LLMs seem like a pretty low-hanging fruit for game studios, even if they don't implement fully immersive player actions upfront. Even just generating background sound or NPC banter during build automatically seems like a no-brainer to me, enabling huge savings on content writing - and that's really the lowest-cost solution imaginable.

Re: Audiobox: Meta's new foundation research model for audio generation

#85

This a fantastic new development in the AI Audio space! However, it's quite disappointing that the model is closed sourced. Nonetheless, Alibaba's equivalent was released earlier in Nov and it's open-sourced! https://github.com/QwenLM/Qwen-Audio Does anyone have suggestions for how to integrate this into your tech stack via an internal API? Interested to hear the varying thoughts on this. From what I softly understan…

Normally I want basically everything to be open source. But as soon as I saw that audio restyling demo, I began to feel concerned they may be open sourcing this. The model can take a sample speakers voice, new text to speak, and also a description of a new location (like a cathedral with many echoes, or other background noises) and produce a convincing new audio sample. This technology will present serious challenges…

this technology is already exploited in the wild for scams. probably for years. it's too late to worry or try stop. the question is what can we do about it. and the first is to make people aware of it.

Re: Audiobox: Meta's new foundation research model for audio generation

#86
The speed of progress is just incredible. I've been using all kinds of different TTS engines for years now and the rapid pace of advancement is awesome. I usually generate all my audiobooks from ebooks and articles and the quality and stability (think artifacts) has gone up so much in the last few months.

Re: Audiobox: Meta's new foundation research model for audio generation

#88
post #6

If I shutdown every voice other than the optimist's one in my head, this, along with other recent AI research, will mark the advent of never-seen-before role play game possibilities. If the current pace of progress continues, we'll see games with complete narrative freedom for players, where you aren't limited to pre-written answers anymore, but can actually talk to in-game characters with your actual voice, goals, a…

> The more rational voices in my mind, though, become more and more afraid of a world where the only thing you can trust is people sitting right in front of you.

You could care less if the content you already consume on forums like HN were generated by AGI. Just sayin.

Re: Audiobox: Meta's new foundation research model for audio generation

#90
post #33
post #31

Earlier quoted context omitted.

This already exists, ex. Audimee: https://audimee.com/

And the millions of other RVC websites. Musicfy, Uberduck, Coversai, Kitsai, FakeYou, Voicemyai, Voicify, Bangerapp, Tryreplay, Weightsgg ... RVC is so easy anyone can spin up a website for it. No moat. Over a hundred thousand trained weights files in the open, so it's easy to bootstrap.

Yeah but I tried Audimee too and they are working on giving best quality voices. It is not about u train or use RVC, it is about quality and save time.
Post reply on HN