Ovi: Twin backbone cross-modal fusion for audio-video generation
81–90 of 122 posts
Re: Ovi: Twin backbone cross-modal fusion for audio-video generation
#82Earlier quoted context omitted.
OpenAI wowed the world with a video model that was also a script writing, editing, dubbing, foley, and music model. Kling still has the best proprietary video model, but Sora 2 is so smart that you don't need to edit anything if your target is social. I don't see how Runway, Pika, or the rest of the purely foundation video model startups survive against the giants and the incredible open source Chinese models. They'v…
There's no moat on the technology, especially with China around. So the only moat is distribution. We're still in the crazy phase of serious changes and if you're betting too deep on one architecture, ah well... https://howlin-wang.github.io/svg/
Re: Ovi: Twin backbone cross-modal fusion for audio-video generation
#83Earlier quoted context omitted.
"Never" is quite short-sighted. Most people would use these tools for personal use, if nothing else. Seeing a celebrity, themselves, their friends, etc., act out any scenario they can think of is quite an appealing proposition. And porn, of course, for better or worse. In the long-term, this has the potential to significantly change how media is created and consumed. Feature films produced by large studios will undou…
Man, am I ever getting tired of replying to the same irrelevant points over and over again. > Most people would use these tools for personal use Not what we're talking about. Not "personalized media", not large studios "leveraging the technology", not "visual effects". See: "blockbuster movies produced by a guy in his basement for <$1000".
If you're unable to draw a line between the points I made and "blockbuster movies produced by a guy in his basement for <$1000", that's on you.
Re: Ovi: Twin backbone cross-modal fusion for audio-video generation
#84Earlier quoted context omitted.
Never. I've seen people instantly go from liking a static image to disliking it upon learning it was AI generated. The same applies to other kinds of media. No matter how "good" it is, knowing that it was created by an unfeeling algorithm ruins it for most people.
I read a while ago that big scientific ideas take about 50 years to be accepted. Which basically means they are never accepted. The people who disagree just get old and die. Younger generation who grow up with AI will just think it’s normal, like we think being connected to the internet via a rectangle you keep in your pocket is normal.
Re: Ovi: Twin backbone cross-modal fusion for audio-video generation
#85How long until we see blockbuster movies produced by a guy in his basement for <$1000?
Not 1000 but Star Wreck, iron sky and Kung Fury are pretty good. https://www.energiavfx.com https://m.youtube.com/watch?v=bS5P_LAqiVg Im sure more wil follow.
(loud music warning)
Re: Ovi: Twin backbone cross-modal fusion for audio-video generation
#86Earlier quoted context omitted.
I’d say it’s the other way around - it took 50 years for EVEN A SCIENTIFIC IDEA - with proof to be accepted. That should have happened super quick. But it didn’t. My point is that you and I will probably never accept it - but our kids will never even think it’s weird in the first place.
That's not the other way around, that's my point. A scientific idea will eventually be accepted because its objective truth makes it inevitable in spite of resistance. Wide acceptance of AI movies is no more inevitable than wide acceptance of bellbottom jeans--it's simply a matter of like or dislike. From what I've seen, people have a strong aversion to it and no particular reason to overcome that aversion. So far no…
It's inevitable because you won't be able to tell the difference.
Re: Ovi: Twin backbone cross-modal fusion for audio-video generation
#87Earlier quoted context omitted.
This is defo true until the moment it gets so good people can't tell.
Gonna be pretty hard to pass off a whole movie as real when none of the "actors" exist.
Re: Ovi: Twin backbone cross-modal fusion for audio-video generation
#88Earlier quoted context omitted.
The other option is to rent a 5090 in the cloud. Probably less than 0.50 per hour at most providers.
If people are interested, I could split up each of our MI300x into 4 and then charge $0.50 (1/4th our current rate). You'd get 48GB of vram instead of 32GB and it would be HBM3 instead of GDDR7 (5.3TB/s vs 1.7TB/s). The only catch is that I'd need to get 32 people who want VMs like this since I would have to do it for the entire box of compute. Wan2.2 runs just fine on AMD.
Though only a shared A40/A100 are in that price range.
Re: Ovi: Twin backbone cross-modal fusion for audio-video generation
#89Earlier quoted context omitted.
Never. I've seen people instantly go from liking a static image to disliking it upon learning it was AI generated. The same applies to other kinds of media. No matter how "good" it is, knowing that it was created by an unfeeling algorithm ruins it for most people.
This is defo true until the moment it gets so good people can't tell.
they're not doing enough to optimize AI data generators for dopamine release with animalistic obsession. Instead they focus on scientific indistinguishablilitiness, and people aren't liking that. IMO that's has been an ongoing and growing costly mistake.
Re: Ovi: Twin backbone cross-modal fusion for audio-video generation
#90Also this model seems to benefit noticeably from having both Cuda >= 12.8 and Torch >= 2.8, and separately SageAttention over Flash 2. But I have yet to see any cache threshold with Easy or Tea that doesn’t get a bit postmodern.