Live data from Hacker News

Veo 2: Our video generation model

deepmind.google

91–100 of 342 posts

Re: Veo 2: Our video generation model

#91

This might be a dumb question to ask, but what exactly is this useful for? B-Roll for YouTube videos? I'm not sure why so much effort is being put into something like this when the applications are so limited.

If you want to train a model to have a general understanding of the physical world, one way is to show it videos and ask it to predict what comes next, and then evaluate it on how close it was to what actually came next.

To really do well on this task, the model basically has to understand physics, and human anatomy, and all sorts of cultural things. So you're forcing the model to learn all these things about the world, but it's relatively easy to train because you can just collect a lot of videos and show the model parts of them -- you know what the next frame is, but the model doesn't.

Along the way, this also creates a video generation model - but you can think of this as more of a nice side effect rather than the ultimate goal.

Re: Veo 2: Our video generation model

#92
post #31

Earlier quoted context omitted.

Everyone has access to YouTube. It’s safe to assume that Sora was trained on it as well.

All you can eat? Surely they charge a lot for that, at least. And how would you even find all the videos?

Who says they've talked to Google about it at all?

I can't speak to OpenAI but ByteDance isn't waiting for permission.

Re: Veo 2: Our video generation model

#93
post #48

OpenAI is like the super luxurious yacht all pretty and shiny, while Google's AI department is the humongous nuclear submarine at least 5 times bigger than the yacht with a relatively cool conning tower, but not that spectacular to look at. Like the tanker which is still steering to fully align with the course people expect it to be, which they don't recognize that it will soon be there and be capable of rolling over…

Or, using Occams Razor; Sundar is a shit CEO and is playing catchup with a company largely fueled by innovations created at Google but never brought to market because it would eat into ads revenue.

That, or they have a secret super human intelligence under wraps at the pentagon.

Re: Veo 2: Our video generation model

#94

I appreciate they posted the skateboarding video. Wildly unrealistic whenever he performs a trick - just morphing body parts. Some of the videos look incredibly believable though.

Cracks in the system are often places where artists find the new and interesting. The leg swapping of the ice skater is mesmerizing in its own way. It would be useful to be able to direct the models in those directions.

Re: Veo 2: Our video generation model

#95

I appreciate they posted the skateboarding video. Wildly unrealistic whenever he performs a trick - just morphing body parts. Some of the videos look incredibly believable though.

Just pretend it's a movie about a shape shifter alien and it's just trying it's best at ice skating, art is subjective like that doesn't it? I bet Salvador Dali would have found those morphing body parts highly amusing.

Re: Veo 2: Our video generation model

#96
post #27

Earlier quoted context omitted.

> I'm not surprised this isn't open to the public by Google yet, Closed models aren't going to matter in the long run. Hunyuan and LTX both run on consumer hardware and produce videos similar in quality to Sora Turbo, yet you can train them and prompt them on anything. They fit into the open source ecosystem which makes building plugins and controls super easy. Video is going to play out in a way that resembles image…

> Hunyuan and LTX both run on consumer hardware Are there other versions than the official? > An NVIDIA GPU with CUDA support is required. > Recommended: We recommend using a GPU with 80GB of memory for better generation quality. https://github.com/Tencent/HunyuanVideo > I am getting CUDA out of memory on an Nvidia L4 with 24 GB of VRAM, even after using the bfloat16 optimization. https://github.com/Lightricks/LTX-Vi…

Yes you can, with some limitations

https://github.com/Tencent/HunyuanVideo/issues/109

Re: Veo 2: Our video generation model

#97

This might be a dumb question to ask, but what exactly is this useful for? B-Roll for YouTube videos? I'm not sure why so much effort is being put into something like this when the applications are so limited.

We're preparing to use video generation (specifically image+text => video so we can also include an initial screenshot of the current game state for style control) for generating in-game cutscenes at our video game studio. Specifically, we're generating them at play-time in a sandbox-like game where the game plays differently each time, and therefore we don't want to prerecord any cutscenes.

Okay, so is the aim to run this locally on a client's computer or served from a cloud? How does the math work out where it's not just easier at that point to render it in game?

Re: Veo 2: Our video generation model

#100

This might be a dumb question to ask, but what exactly is this useful for? B-Roll for YouTube videos? I'm not sure why so much effort is being put into something like this when the applications are so limited.

Back when computers took up a whole room, you'd also have asked: "but what exactly is this useful for? B-Roll some simple calculations that anybody can do with a piece of paper and a pen."? Think 5-10 years into the future, this is a stepping stone

this is kind of an unfair comparison. Whats the endpoint of generating AI videos? What can this do that is useful, contributes something to society, has artistic value, etc etc. We can make educational videos with a script but its also pretty easy for motivated parties to do that already, and its getting easier as cameras get better and smaller. I think asking "whats the point of this" is at least fair.
Post reply on HN