Live data from Hacker News

Veo 2: Our video generation model

deepmind.google

201–210 of 342 posts

Re: Veo 2: Our video generation model

#202
post #28

Is it just me or do all these models generate everything in a weird pseudo-slow motion framerate?

I mean, I'm not sure it's done deliberately but... if I was trying to guarantee video gen was always 5 seconds in a consistent manner and the gen process was highly non-deterministic then if the resultant video would have only been 3 seconds I'd stretch it out, interpolate the frames, and then send it down the pipes.

Another point to consider is that if my generative video system isn't good at maintaining world consistency, then doing a slow-motion video gives the illusion of a long video while being able to maintain a smaller "world context".

Re: Veo 2: Our video generation model

#203
I've already started not to notice the quality differences in the photo-like images produced by each image generation model.

Now, examples of image or video generation models showing off how great they are should be stickman drawings or stickman videos. As far as I know, no model has been able to do that properly yet. If a model can do it well, it will be a huge breakthrough.

Re: Veo 2: Our video generation model

#204

Earlier quoted context omitted.

What officials actually say doesn't make a difference anymore. People do not get bamboozled because of lack of facts. People who get bamboozled are past facts.

Off topic from the video AI thread, but to elaborate on your point: people believe what they want, based on what they have been primed to believe from mass media. This is mainly the normal TV and paper news, filtered through institutions like government proclamations, schools, and now supercharged by social media. This is why the "narrative" exists, and news media does the consensus messaging of what you should belie…

What are you talking about? News media LOVE twitter/X, it is where they get all their stories from and journalists are notoriously addicted to it, to their detriment.

Re: Veo 2: Our video generation model

#205
post #131

I got access to the preview, here's what it gave me for "A pelican riding a bicycle along a coastal path overlooking a harbor" - this video has all four versions shown: https://static.simonwillison.net/static/2024/pelicans-on-bic... Of the four two were a pelican riding a bicycle. One was a pelican just running along the road, one was a pelican perched on a stationary bicycle, and one had the pelican wearing a weird…

Well yeah, if you look closely at the example videos on the site, one of them is not quite right either:

> Prompt: The sun rises slowly behind a perfectly plated breakfast scene. Thick, golden maple syrup pours in slow motion over a stack of fluffy pancakes, each one releasing a soft, warm steam cloud. A close-up of crispy bacon sizzles, sending tiny embers of golden grease into the air. [...]

In the video, the bacon is unceremoniously slapped onto the pancakes, while the prompt sounds like it was intended to be a separate shot, with the bacon still in the pan? Or, alternatively, everything described in the prompt should have been on the table at the same time?

So, yet again: AI produces impressive results, but it rarely does exactly what you wanted it to do...

Re: Veo 2: Our video generation model

#206
Imho is stunning, yet what is happening there is super dangerous.

These videos will and may be too realistic.

Our society is not prepared for this kind of reality "bending" media. These hyperrealistic videos will be the reason for hate and murder. Evil actors will use it to influence elections on a global scale. Create cults around virtual characters. Deny the rules of physics and human reason. And yet, there is no way for a person to detect instantly that he is watching a generated video. Maybe now, but in 1 year, it will be indistinguishable from a real recorded video

Re: Veo 2: Our video generation model

#207
post #131

I got access to the preview, here's what it gave me for "A pelican riding a bicycle along a coastal path overlooking a harbor" - this video has all four versions shown: https://static.simonwillison.net/static/2024/pelicans-on-bic... Of the four two were a pelican riding a bicycle. One was a pelican just running along the road, one was a pelican perched on a stationary bicycle, and one had the pelican wearing a weird…

Here's one of a penguin paragliding and it's surprisingly realistic https://x.com/Plinz/status/1868885955597549624

Re: Veo 2: Our video generation model

#208
post #131

I got access to the preview, here's what it gave me for "A pelican riding a bicycle along a coastal path overlooking a harbor" - this video has all four versions shown: https://static.simonwillison.net/static/2024/pelicans-on-bic... Of the four two were a pelican riding a bicycle. One was a pelican just running along the road, one was a pelican perched on a stationary bicycle, and one had the pelican wearing a weird…

There's another important contender in the space: Hunyuan model from Tencent My company (Nim) is hosting Hunyuan model, so here's a quick test (first attempt) at "pelican riding a bycicle" via Hunyuan on Nim: https://nim.video/explore/OGs4EM3MIpW8 I think it's as good, if not better than Sora / Veo

> A whimsical pelican, adorned in oversized sunglasses and a vibrant, patterned scarf, gracefully balances on a vintage bicycle, its sleek feathers glistening in the sunlight. As it pedals joyfully down a scenic coastal path, colorful wildflowers sway gently in the breeze, and azure waves crash rhythmically against the shore. The pelican occasionally flaps its wings, adding a playful touch to its enchanting ride. In the distance, a serene sunset bathes the landscape in warm hues, while seagulls glide gracefully overhead, celebrating this delightful and lighthearted adventure of a pelican enjoying a carefree day on two wheels.

What does it produce for “A pelican riding a bicycle along a coastal path overlooking a harbor”?

Or, what do Sora and Veo produce for your verbose prompt?

Re: Veo 2: Our video generation model

#209

This might be a dumb question to ask, but what exactly is this useful for? B-Roll for YouTube videos? I'm not sure why so much effort is being put into something like this when the applications are so limited.

It's got a lot of potential as a way for google to get paid for other people's skills and hard work instead of the people that made all of that "data".

It’s kind of hilarious that anybody considers this “democratizing” creating media. How many people that need a video clip are going to be capable of running an open version of this themselves? The wonky “open” models aren’t even close. How much do you think these services are going to cost once the introductory period financed by race-to-the-bottom money stops? OpenAI already charges $200/mo if you want to be guaranteed more than 30-60 minutes of Advanced Voice. The introductory period exists solely to get people engaged enough to push through blatantly stealing millions of artists creative output so they can have a beautiful tool they sell to Hollywood for a whole lot of money that’s still less than traditional vfx, and to m everyone gets to dink around in the useless free models or too-expensive-for-most prosumer tools and people with expensive video card arrays or the functional equivalent will still be niche tinkering hobbyists with inferior tooling and models and the skilled commercial artists still employed are being paid shit because of market forces. Great job SV. Making the world a better place.

Re: Veo 2: Our video generation model

#210
post #206

Imho is stunning, yet what is happening there is super dangerous. These videos will and may be too realistic. Our society is not prepared for this kind of reality "bending" media. These hyperrealistic videos will be the reason for hate and murder. Evil actors will use it to influence elections on a global scale. Create cults around virtual characters. Deny the rules of physics and human reason. And yet, there is no w…

The society voted with their money. Google refrained from launching their early chatbots and image generation tools due to perceived risks of unsafe and misleading content being generated, and got beaten to the punch in the market. Of course now they'll launch early and often, the market has spoken.
Post reply on HN