Live data from Hacker News

Ovi: Twin backbone cross-modal fusion for audio-video generation

github.com

41–50 of 122 posts

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#41
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

Never. I've seen people instantly go from liking a static image to disliking it upon learning it was AI generated. The same applies to other kinds of media. No matter how "good" it is, knowing that it was created by an unfeeling algorithm ruins it for most people.

New things will be possible that aren’t today. You’ll be able to pick the stars in your movie. The home base can be your childhood home. Your unrequited love can be virtually fulfilled. Etc etc.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#42
post #15
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

I don't know how long it will take for photorealistic movies but I'm looking forward to that Hyperion animated movie I'd make myself because nobody wants or is able to.

honestly i think there is a future in this.. "hey AI model, generate a movie from this book" may indeed be better than letting a director mangle it into something meant to sell tickets. most book -> movie adaptations are not setting a high bar

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#43
post #41

Earlier quoted context omitted.

Never. I've seen people instantly go from liking a static image to disliking it upon learning it was AI generated. The same applies to other kinds of media. No matter how "good" it is, knowing that it was created by an unfeeling algorithm ruins it for most people.

New things will be possible that aren’t today. You’ll be able to pick the stars in your movie. The home base can be your childhood home. Your unrequited love can be virtually fulfilled. Etc etc.

A "blockbuster movie" implies commercial use. I think actual stars would object to their likeness being used in that way.

If you're talking about people firing up the ol' 5090 to make a "movie" about their favorite streamer falling madly in love with them for, ahem, personal use, I have no doubt that people will do that. And I will do everything in my power to avoid associating with such brain-rotted cretins.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#44

Earlier quoted context omitted.

To be fair, a lot of movies in theaters have uncanny valleys. A lot of CGI feels that way to me.

The limits of CGI have gotten so good that if you are noticing the CGI, it's b/c someone skimped on the budget for the particular movie where you saw the uncanny valley. (Of course, excluding the obvious "that guy just knocked down a building!" CGI)

Yeah that's the thing, with good execution of a plausible effect you don't even realize you're looking at CGI. It's the toupée fallacy.

Obligatory Jonas Ussing plug: https://www.youtube.com/watch?v=7ttG90raCNo&list=PLgdTaHO8FL...

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#45
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

Besides people with weird fetishes, who actually enjoys looking at AI "art"?

Photos, movies, animation, recorded music, computer games, CGI props in movies, CGI characters in movies were all denounced as "not real art", until good counter-examples appeared.

AI is but a tool; if there is an artist using them, real art can be created, as with any other tool.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#46
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

I think "blockbuster movie" is a moving target, so it's a bit hard to know

It's a relatively well defined measure of success though: a movie which is popular and high-grossing.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#47
post #13
post #6

Seems like the video model is based on Wan2.2. Lots of activity around Wan lately. It’s nice to see flexible open models make a strong showing against the massively funded closed competitors like OpenAI and Runway.

And Google.

[dead]

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#49
post #17

Heh, I used to work for Nokia's Ovi - basically, gsuite for nokia phones (my group did map search) - the official explanation was "Ovi is Finnish for Door", the internal joke was "Ovi is Hungarian for Kindergarten". I couldn't find any backstory about the name here, though.

[dead]

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#50
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

I've yet to see continuity working. Small clips is one thing. To have the same character, wearing the same clothes, revisit environments, with the same lighting and post processing is very different. I think we'll see AGI first.

That and control, there's an enormous difference between letting a model just YOLO a prompt, filling in 95% of the details with vaguely plausible whatever, and getting a model to execute on a meticulously planned out shot with every last detail in order.
Post reply on HN