Live data from Hacker News

Ovi: Twin backbone cross-modal fusion for audio-video generation

github.com

51–60 of 122 posts

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#52
post #7

Kinda terrifying. And it can run in 32GB of VRAM? Anyone with a 5090 can start spewing out believable fake videos.

Yeah and even with the server it’s really cheap - things like omnihuman are I think better but MUCH more expensive to run

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#53
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

Besides people with weird fetishes, who actually enjoys looking at AI "art"?

I enjoyed just about everything in here: https://andymasley.substack.com/p/a-ton-of-ai-images-ive-mad...

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#54

Earlier quoted context omitted.

I read a while ago that big scientific ideas take about 50 years to be accepted. Which basically means they are never accepted. The people who disagree just get old and die. Younger generation who grow up with AI will just think it’s normal, like we think being connected to the internet via a rectangle you keep in your pocket is normal.

Scientific ideas have the benefit of being objectively true. AI movies are not a "scientific idea". Liking them is a matter of taste, and there are plenty of things that never catch on.

I’d say it’s the other way around - it took 50 years for EVEN A SCIENTIFIC IDEA - with proof to be accepted. That should have happened super quick. But it didn’t.

My point is that you and I will probably never accept it - but our kids will never even think it’s weird in the first place.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#55
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

> How long until we see blockbuster movies produced by a guy in his basement for Probably never. If AI is good enough to cover all the skills needed to do what would currently make a blockbuster movie for less than $1000, the demand for movies will be small enough relative to supply that there will be no such thing as a “blockbuster movie”

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#56
post #5
post #4

Lazyweb: Are these related? If so, how? https://news.ycombinator.com/item?id=45603435 https://news.ycombinator.com/item?id=45652726

When a new open weights AI model comes out, opportunists register a domain using its name and start hosting it hoping to make a buck with SEO. Easier than ever now, as AI-assisted coding tools will build you that generic landing page and basic UI.

If they're actually hosting it and providing a service, that wouldn't be a bad thing. Ovi's Apache license allows commercial use. They would only be infringing on the license by not publishing it and the original copyright, and would be morally bankrupt by not disclosing the original project and passing it off as their own.

But I also suspect that most of these are indeed SEO scammers, that there's no actual service, and that all payments are pocketed. It might take a few days for the scam to be reported and the site taken down, but it's likely enough to get a few hundred bucks out of it. They'll never be pursued because of where they live, and they can have many of these up in no time, thanks to AI, as you say.

What a sad state of affairs that no "AI" company or government is taking seriously.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#57
post #41

Earlier quoted context omitted.

Never. I've seen people instantly go from liking a static image to disliking it upon learning it was AI generated. The same applies to other kinds of media. No matter how "good" it is, knowing that it was created by an unfeeling algorithm ruins it for most people.

New things will be possible that aren’t today. You’ll be able to pick the stars in your movie. The home base can be your childhood home. Your unrequited love can be virtually fulfilled. Etc etc.

Yes, and none of these hyperpersonal movies will be a blockbuster; they'll be lucky to have audiences requiring the fingers of more than one hand to count, because everyone will have their own hyperpersonal preferences.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#58
post #41

Earlier quoted context omitted.

New things will be possible that aren’t today. You’ll be able to pick the stars in your movie. The home base can be your childhood home. Your unrequited love can be virtually fulfilled. Etc etc.

Yes, and none of these hyperpersonal movies will be a blockbuster; they'll be lucky to have audiences requiring the fingers of more than one hand to count, because everyone will have their own hyperpersonal preferences.

A good thing, since most such movies will be made to keep one hand occupied.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#59

At this rate, in a few months we will have probably some high quality shorts entirely generated by this.

It's funny you mention this, I was just thinking this other day we may eventually be in a future where a group hangout party could look like this:

1. Goes to friends' place 2. Usual drinks, whatever gets you going activity 3. Each person writes a prompt 4. Chain them together 5. Watch the resulting movie together

That sounds hilarious and I can't wait to try

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#60
post #8

How long until we see blockbuster movies produced by a guy in his basement for <$1000?

It depends on what you mean by "blockbuster movie", as even those can have awful visual effects. But a short film released a few months ago made entirely with video generation tools was surprisingly decent[1]. It still requires a talented "director" to have the right vision to guide the project, but the tools are there.

Before we see this and higher level of quality accessible to enthusiasts, we'll see these tools adopted by mainstream studios first, which is starting to happen.

I'm a firm "AI" skeptic, but if this technology has revolutionized anything, it has been image generation. A few years ago it was science fiction to have the quality of upscaling we take for granted today. I reckon the same will happen with video generation as well a few years from now. Unlike "ASI" and "AGI", these improvements are achievable with better engineering, and don't necessarily require a breakthrough.

[1]: https://news.ycombinator.com/item?id=44564697

Post reply on HN