Live data from Hacker News

Ovi: Twin backbone cross-modal fusion for audio-video generation

github.com

31–40 of 122 posts

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#31

Earlier quoted context omitted.

Never. I've seen people instantly go from liking a static image to disliking it upon learning it was AI generated. The same applies to other kinds of media. No matter how "good" it is, knowing that it was created by an unfeeling algorithm ruins it for most people.

I read a while ago that big scientific ideas take about 50 years to be accepted. Which basically means they are never accepted. The people who disagree just get old and die. Younger generation who grow up with AI will just think it’s normal, like we think being connected to the internet via a rectangle you keep in your pocket is normal.

Scientific ideas have the benefit of being objectively true.

AI movies are not a "scientific idea". Liking them is a matter of taste, and there are plenty of things that never catch on.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#33

Earlier quoted context omitted.

Never is a long time. The resistance will erode.

"People will like AI movies because it's inevitable" Circular reasoning. If you can't answer WHY people should come to like AI movies, then you have nothing to say.

People will like them when they’re good content. Right now we’re stuck in the in between where it’s kind of all or nothing ai, but it will get grayer when the feedback loops are tighter and building ai movies is more interactive. Same thing with any special effects really

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#34
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

I fully expect that we will see an AI video project that gets to Skibidi Toilet levels of cultural reach within the next two years, but 'blockbluster' implies a level of financial success that is much harder to predict.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#35
post #3

mindblowing - but still in the uncanny valley. and I guess it's cute that many of the characters live in a world where AI has caused an apocolypse, but is that really the message they want to lead with?

> but still in the uncanny valley

that and the guitar player behind the singer in the concert example has three arms :)

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#36
post #13
post #6

Seems like the video model is based on Wan2.2. Lots of activity around Wan lately. It’s nice to see flexible open models make a strong showing against the massively funded closed competitors like OpenAI and Runway.

And Google.

OpenAI wowed the world with a video model that was also a script writing, editing, dubbing, foley, and music model.

Kling still has the best proprietary video model, but Sora 2 is so smart that you don't need to edit anything if your target is social.

I don't see how Runway, Pika, or the rest of the purely foundation video model startups survive against the giants and the incredible open source Chinese models. They've got to be sweating bullets right now.

Everyone's also sleeping on xAI's high quality and insanely fast video model (10 second generations) that they're giving away completely for free without watermarks.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#37
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

I've yet to see continuity working. Small clips is one thing. To have the same character, wearing the same clothes, revisit environments, with the same lighting and post processing is very different.

I think we'll see AGI first.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#38

Earlier quoted context omitted.

"People will like AI movies because it's inevitable" Circular reasoning. If you can't answer WHY people should come to like AI movies, then you have nothing to say.

People will like them when they’re good content. Right now we’re stuck in the in between where it’s kind of all or nothing ai, but it will get grayer when the feedback loops are tighter and building ai movies is more interactive. Same thing with any special effects really

I think AI has its place in special effects, but "making a blockbuster movie for $1000" requires replacing all the actors, music, cinematography, everything that makes a movie "art" except maybe the plot. And I have never seen anyone respond well to AI "art" of any form. I've seen some fairly passable (if a bit boring) AI music passed around on Reddit and it was universally met with disgust simply because it was AI.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#39
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

The main constraint to good movies is good actors and good scripts. So presumably an AI would have to write well or perform well to do that.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#40
post #6

Seems like the video model is based on Wan2.2. Lots of activity around Wan lately. It’s nice to see flexible open models make a strong showing against the massively funded closed competitors like OpenAI and Runway.

Discussed here:

Wan – Open-source alternative to VEO 3 - https://news.ycombinator.com/item?id=44928997 - Aug 2025 (38 comments)

Post reply on HN