Live data from Hacker News

Veo 2: Our video generation model

deepmind.google

131–140 of 342 posts

Re: Veo 2: Our video generation model

#131
I got access to the preview, here's what it gave me for "A pelican riding a bicycle along a coastal path overlooking a harbor" - this video has all four versions shown:

https://static.simonwillison.net/static/2024/pelicans-on-bic...

Of the four two were a pelican riding a bicycle. One was a pelican just running along the road, one was a pelican perched on a stationary bicycle, and one had the pelican wearing a weird sort of pelican bicycle helmet.

All four were better than what I got from Sora: https://simonwillison.net/2024/Dec/9/sora/

Re: Veo 2: Our video generation model

#132
post #131

I got access to the preview, here's what it gave me for "A pelican riding a bicycle along a coastal path overlooking a harbor" - this video has all four versions shown: https://static.simonwillison.net/static/2024/pelicans-on-bic... Of the four two were a pelican riding a bicycle. One was a pelican just running along the road, one was a pelican perched on a stationary bicycle, and one had the pelican wearing a weird…

His little bike helmet is adorable

Re: Veo 2: Our video generation model

#133

It's interesting they host these videos on YouTube, cause it signals they're fine with AI generated content. I wonder if Google forgets that the creators themselves are what makes YouTube interesting for viewers.

What makes you think that viewers wouldn't be watching AI generated content? Considering the possibilities of fake videos, I'm sure that it can be very engaging. And the costs are zero.

Re: Veo 2: Our video generation model

#134
post #131

I got access to the preview, here's what it gave me for "A pelican riding a bicycle along a coastal path overlooking a harbor" - this video has all four versions shown: https://static.simonwillison.net/static/2024/pelicans-on-bic... Of the four two were a pelican riding a bicycle. One was a pelican just running along the road, one was a pelican perched on a stationary bicycle, and one had the pelican wearing a weird…

As long as at least one option is exactly what you asked for throwing variations at you that don't conform to 100% of your prompt seems like it could be useful if it gives the model leeway to improve the output in other aspects.

Re: Veo 2: Our video generation model

#135
post #7

Winning 2:1 in user preference versus sora turbo is impressive. It seems to have very similar limitations to sora. For example- the leg swapping in the ice skating video and the bee keeper picking up the jar is at a very unnatural acceleration (like it pops up). Though by my eye maybe slightly better emulating natural movement and physics in comparison to sora. The blog post has slightly more info: >at resolutions up…

It looks Sora is actually the worst performer in the benchmarks, with Kling being the best and others not far behind. Anyways, I strongly suspect that the funny meme content that seems to be the practical uses case of these video generators won't be possible on either Veo or Sora, because of copyright, PC, containing famous people, or other 'safety' related reasons.

I’ve been using Kling a lot recently and been really impressed, especially by 1.5.

I was so excited to see Sora out - only to see it has most of the same problems. And Kling seems to do better in a lot of benchmarks.

I can’t quite make sense of it - what OpenAI were showing when they first launched Sora was so amazing. Was it cherry picked? Or was it using loads more compute than what they’ve release?

Re: Veo 2: Our video generation model

#136
post #91

Earlier quoted context omitted.

If you want to train a model to have a general understanding of the physical world, one way is to show it videos and ask it to predict what comes next, and then evaluate it on how close it was to what actually came next. To really do well on this task, the model basically has to understand physics, and human anatomy, and all sorts of cultural things. So you're forcing the model to learn all these things about the wor…

It doesn’t have to understand anything, none of these demonstrate reasoning or understanding. All these models have just “seen” enough videos of all those things to build a probability distribution to predict the next step. This is not bad, or make it inherently dumb, a major component of human intelligence is built on similar strategies. I couldn’t tell what grammatical rules are broken in text or what physical rule…

That's like saying that your brain doesn't understand anything, it just analyzes the visual data coming in via your eyes and predicts the next step of reality

Re: Veo 2: Our video generation model

#137

I appreciate they posted the skateboarding video. Wildly unrealistic whenever he performs a trick - just morphing body parts. Some of the videos look incredibly believable though.

The honey, Peruvian women, swimming dog, bee keeper, DJ etc. are stunning. They’re short but I can barely find any artifacts.

The prompt for the honey video mentions ending with a shot of an orange. The orange just...isn't there, though?

Re: Veo 2: Our video generation model

#138

Earlier quoted context omitted.

It looks Sora is actually the worst performer in the benchmarks, with Kling being the best and others not far behind. Anyways, I strongly suspect that the funny meme content that seems to be the practical uses case of these video generators won't be possible on either Veo or Sora, because of copyright, PC, containing famous people, or other 'safety' related reasons.

I’ve been using Kling a lot recently and been really impressed, especially by 1.5. I was so excited to see Sora out - only to see it has most of the same problems. And Kling seems to do better in a lot of benchmarks. I can’t quite make sense of it - what OpenAI were showing when they first launched Sora was so amazing. Was it cherry picked? Or was it using loads more compute than what they’ve release?

The SORA model available to the public is a smaller, distilled model called SORA Turbo. What was originally shown was a more capable model that was probably too slow to meet their UX requirements for the sora.com user interface.

Re: Veo 2: Our video generation model

#139

Earlier quoted context omitted.

“Won” what exactly? I have no issues running stable diffusion locally. Since Llama3.3 came out it is my first stop for coding questions, and I’m only using closed models when llama3.3 has trouble. I think it’s fairly clear that between open weights and LLMs plateauing, the game will be who can build what on top of largely equivalent base models.

The quality for SD is no where near the clear leaders.

> The quality for SD is no where near the clear leaders.

It absolutely is. Moreover, the tools built on top of SD (and now Flux) are superior to any commercial vertical.

The second-place companies and research labs will continue to release their models as open source, which will cause further atrophy to the value of building a foundation model. Value will accrue in the product, as has always been the case.

Re: Veo 2: Our video generation model

#140
post #131

I got access to the preview, here's what it gave me for "A pelican riding a bicycle along a coastal path overlooking a harbor" - this video has all four versions shown: https://static.simonwillison.net/static/2024/pelicans-on-bic... Of the four two were a pelican riding a bicycle. One was a pelican just running along the road, one was a pelican perched on a stationary bicycle, and one had the pelican wearing a weird…

His little bike helmet is adorable

The AI safety team was really proud of that one.
Post reply on HN