Live data from Hacker News

Ovi: Twin backbone cross-modal fusion for audio-video generation

github.com

71–80 of 122 posts

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#71

Earlier quoted context omitted.

Scientific ideas have the benefit of being objectively true. AI movies are not a "scientific idea". Liking them is a matter of taste, and there are plenty of things that never catch on.

I’d say it’s the other way around - it took 50 years for EVEN A SCIENTIFIC IDEA - with proof to be accepted. That should have happened super quick. But it didn’t. My point is that you and I will probably never accept it - but our kids will never even think it’s weird in the first place.

>I’d say it’s the other way around - it took 50 years for EVEN A SCIENTIFIC IDEA - with proof to be accepted.

I recall eerily similar things said about Google Glass..

Maybe AI generation will be used in popular media more often, but purely AI generated content or AI brain rot seems to only appeal to a small crowd of people right now, and I don't see that crowd growing significantly.

Maybe it's a technology problem, as Google Glass was, but I think that's inseparable from the content it actually generates at this non-AGI stage.

Regardless, it sounds very uncertain and perhaps even unlikely that what we see being created now is the future.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#72

At this rate, in a few months we will have probably some high quality shorts entirely generated by this.

It's funny you mention this, I was just thinking this other day we may eventually be in a future where a group hangout party could look like this: 1. Goes to friends' place 2. Usual drinks, whatever gets you going activity 3. Each person writes a prompt 4. Chain them together 5. Watch the resulting movie together That sounds hilarious and I can't wait to try

I'm vaguely reminded of the excellent Jackbox game Tee Fury, in which players submit slogans for T shirts and "art" separately. Players then get to choose from a few options for slogans and designs to make T shirts which are voted on by the group.

I have fond memories of laughing until I was in tears when playing with a group of friends over drinks during the lockdowns in 2020. Something about the process just naturally results in hilarity (especially if you're in a group where you can be offensive).

It's like exquisite corpse for t-shirts. Or, in your case, shorts.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#73

Earlier quoted context omitted.

It's funny you mention this, I was just thinking this other day we may eventually be in a future where a group hangout party could look like this: 1. Goes to friends' place 2. Usual drinks, whatever gets you going activity 3. Each person writes a prompt 4. Chain them together 5. Watch the resulting movie together That sounds hilarious and I can't wait to try

I'm vaguely reminded of the excellent Jackbox game Tee Fury, in which players submit slogans for T shirts and "art" separately. Players then get to choose from a few options for slogans and designs to make T shirts which are voted on by the group. I have fond memories of laughing until I was in tears when playing with a group of friends over drinks during the lockdowns in 2020. Something about the process just natura…

T shirt game is the best jackbox game!

Whenever one of my friend groups is gathered we always make it a point to do an exquisite corpse story on a piece of paper while we’re inebriated in some way xD Video version will be wild

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#74
post #70

Earlier quoted context omitted.

> How long until we see blockbuster movies produced by a guy in his basement for Probably never. If AI is good enough to cover all the skills needed to do what would currently make a blockbuster movie for less than $1000, the demand for movies will be small enough relative to supply that there will be no such thing as a “blockbuster movie”

The same argument could reasonably be used to explain why no YouTube influencer would ever get more than 1,000 subscribers - if everyone can upload videos that anyone can watch, nobody will really be famous because fame will become very evenly distributed, right?

> The same argument could reasonably be used to explain why no YouTube influencer would ever get more than 1,000 subscribers -

No, that would require a radically different argument, in pretty much every way.

> if everyone can upload videos that anyone can watch, nobody will really be famous because fame will become very evenly distributed, right?

No, Youtube makes distribution cheap, but it doesn't substitute for most of the other things that differentiate between videos; most of the skills that provide variation between videos are still there, and not cheaply substituted via YouTube.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#75
This is really amazing, I've been working with AI generation for months and it's amazing how fast separate tools are coming together into one and are usable on your own local machine.

I've been using Ovi for about a week and it's a blast. Like all AI gen, it's a slot machine and even putting in good inputs might lead to bad outputs, but if you run it enough you'll get something good or usable.

I've definitely made many things that look and sound real with both I2V and T2V, albeit T2V tends to look more like 90s tv quality at times, but that also makes it seem more real. If you use Flux SPRO as the image source you can get some pretty realistic looking videos.

I do have a 5090, so it takes about 4 to 5 minutes to make a 5 second clip.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#76
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

Not 1000 but Star Wreck, iron sky and Kung Fury are pretty good.

https://www.energiavfx.com

https://m.youtube.com/watch?v=bS5P_LAqiVg

Im sure more wil follow.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#77
post #7

Kinda terrifying. And it can run in 32GB of VRAM? Anyone with a 5090 can start spewing out believable fake videos.

The other option is to rent a 5090 in the cloud. Probably less than 0.50 per hour at most providers.

If people are interested, I could split up each of our MI300x into 4 and then charge $0.50 (1/4th our current rate). You'd get 48GB of vram instead of 32GB and it would be HBM3 instead of GDDR7 (5.3TB/s vs 1.7TB/s).

The only catch is that I'd need to get 32 people who want VMs like this since I would have to do it for the entire box of compute.

Wan2.2 runs just fine on AMD.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#78
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

> How long until we see blockbuster movies produced by a guy in his basement for Probably never. If AI is good enough to cover all the skills needed to do what would currently make a blockbuster movie for less than $1000, the demand for movies will be small enough relative to supply that there will be no such thing as a “blockbuster movie”

Just like with games visual are only part of the formula. In theory you can make a truly fantastic movie (or game) for next to nothing. I didnt believe this before The man from earth. That cost 200K but it didnt have to.

Edit: perhaps 12 angry men was good enough at the time.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#79
post #36
post #13

Earlier quoted context omitted.

And Google.

OpenAI wowed the world with a video model that was also a script writing, editing, dubbing, foley, and music model. Kling still has the best proprietary video model, but Sora 2 is so smart that you don't need to edit anything if your target is social. I don't see how Runway, Pika, or the rest of the purely foundation video model startups survive against the giants and the incredible open source Chinese models. They'v…

[deleted]

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#80
post #36
post #13

Earlier quoted context omitted.

And Google.

OpenAI wowed the world with a video model that was also a script writing, editing, dubbing, foley, and music model. Kling still has the best proprietary video model, but Sora 2 is so smart that you don't need to edit anything if your target is social. I don't see how Runway, Pika, or the rest of the purely foundation video model startups survive against the giants and the incredible open source Chinese models. They'v…

There's no moat on the technology, especially with China around. So the only moat is distribution. We're still in the crazy phase of serious changes and if you're betting too deep on one architecture, ah well... https://howlin-wang.github.io/svg/
Post reply on HN