Live data from Hacker News

Veo 2: Our video generation model

deepmind.google

151–160 of 342 posts

Re: Veo 2: Our video generation model

#151

This might be a dumb question to ask, but what exactly is this useful for? B-Roll for YouTube videos? I'm not sure why so much effort is being put into something like this when the applications are so limited.

You and your friends gather around the TV to watch a video about the time that you all traveled abroad and met a mysterious stranger. In the film, you witness each other take incredible risks, have intimate private conversations, and change in profound ways. Of course none of it actually happened; your voices and likenesses were fed into the movie generator. And did I mention in the film you’re driving expensive cars and wearing designer clothes?

Re: Veo 2: Our video generation model

#152
Superficially impressive but what is the actual use case of the present state of the art? It makes 10-second demos, fine. But can a producer get a second shot of the same scene and the same characters, with visual continuity? Or a third, etc? In other words, can it be used to create a coherent movie --even a 60-second commercial -- with multiple shots having continuity of faces, backgrounds, and lighting?

This quote suggests not: "maintaining complete consistency throughout complex scenes or those with complex motion, remains a challenge."

Re: Veo 2: Our video generation model

#153

Earlier quoted context omitted.

It doesn’t have to understand anything, none of these demonstrate reasoning or understanding. All these models have just “seen” enough videos of all those things to build a probability distribution to predict the next step. This is not bad, or make it inherently dumb, a major component of human intelligence is built on similar strategies. I couldn’t tell what grammatical rules are broken in text or what physical rule…

That's like saying that your brain doesn't understand anything, it just analyzes the visual data coming in via your eyes and predicts the next step of reality

The brain also does that . It doesn’t do it exclusively, but we do it an awful lot .

we do extensive amount of pattern matching and drop enormous amount of sensory input very quickly because we expect patterns and assume a lot about our surroundings.

Unlearning this is a hard skill to pick up. There are many versions of training from martial arts to meditation that attempt to achieve this .

Point is that alone is not sufficient, the other core component is reasoning and understanding , transformers and learning on data is insufficient .

Parrot and few other animals can imitate human speech very well , that doesn’t mean they are understanding the speech or constructing .

Don’t get me wrong, i am not saying it is not useful, it is , but this attribution of reasoning and understanding to models that foundationally has no such building block is just being impressed by a speaking parrot

Re: Veo 2: Our video generation model

#154

Earlier quoted context omitted.

Back when computers took up a whole room, you'd also have asked: "but what exactly is this useful for? B-Roll some simple calculations that anybody can do with a piece of paper and a pen."? Think 5-10 years into the future, this is a stepping stone

That's comparing apples to oranges though isn't it? Generating videos is the output of the technology, not the tech itself. It would be like someone asking "this computer that takes up a whole room printed out ascii art, what is this useful for?"

all the "creative" gen ai does a thing worse and more annoying than what exists now. the first computers did calculations faster and faster with immediate utility (for defense mostly)

Re: Veo 2: Our video generation model

#155
post #48

OpenAI is like the super luxurious yacht all pretty and shiny, while Google's AI department is the humongous nuclear submarine at least 5 times bigger than the yacht with a relatively cool conning tower, but not that spectacular to look at. Like the tanker which is still steering to fully align with the course people expect it to be, which they don't recognize that it will soon be there and be capable of rolling over…

Or, using Occams Razor; Sundar is a shit CEO and is playing catchup with a company largely fueled by innovations created at Google but never brought to market because it would eat into ads revenue. That, or they have a secret super human intelligence under wraps at the pentagon.

That's the conventional take, but (as far as I can tell), the TPU program was also started under Sundar, which would have been a bold investment at the time, and looks like absolute genius in retrospect.

OpenAI might be well-capitalized, but they're (1) bleeding money, (2) no clear path to profitability, and (3) competing head-to-head with a behemoth who can profitably provide a similar offering at 10-20x cheaper (literally).

Google might be slow out the blocks, but it's not like they've been sitting on their hands for the past decade.

Re: Veo 2: Our video generation model

#156
post #99

just to remind everyone that state of the art was Will Smith Eating Spaghetti in April of 2023 https://arstechnica.com/information-technology/2023/03/yes-v... We're not even done with 2024. Just imagine what's waiting for us in 2025.

But it's the same thing just at a higher fidelity. Which is impressive don't get me wrong. But they are also kinda bad looking. Like even there good examples have so many issues. I just don't see how this gets extrapolated into the ideas in various posts like full length movies, custom TV shows and holodecks or whatever else people dream up. Do we have any examples of tech that just kept improving at exponential or linear rates? Why is everyone so confident it will just keep getting better?

Re: Veo 2: Our video generation model

#157
post #152

Superficially impressive but what is the actual use case of the present state of the art? It makes 10-second demos, fine. But can a producer get a second shot of the same scene and the same characters, with visual continuity? Or a third, etc? In other words, can it be used to create a coherent movie --even a 60-second commercial -- with multiple shots having continuity of faces, backgrounds, and lighting? This quote…

B-roll for YouTube videos.

Re: Veo 2: Our video generation model

#158

Impressive but the page crashed chrome on my iPad!

Might be time for a new iPad. My old-school iPad Air has 2gb of memory and is an absolute hog when loading content-heavy websites.

It crashed Safari on iPhone 16 Pro Max, I doubt it's the device.

The website is horrible on resources.

Re: Veo 2: Our video generation model

#160
post #65

Earlier quoted context omitted.

I mean, I have trained myself on Youtube. Why can't a silicon being train itself on Youtube as well?

Because silicon is a robot. A camcorder can't catch a flick with me in the theater even if I dress it up like a muppet.

Not with that attitude.

A corporation "is a person" with all the rights that come along with that - free speech etc.

Post reply on HN