Veo 2: Our video generation model
121–130 of 342 posts
Re: Veo 2: Our video generation model
#122This looks great, but I'm confused by this part: > Veo sample duration is 8s, VideoGen’s sample duration is 10s, and other models' durations are 5s. We show the full video duration to raters. Could the positive result for Veo 2 mean the raters like longer videos? Why not trim Veo 2's output to 5s for a better controlled test? I'm not surprised this isn't open to the public by Google yet, there's a huge amount of volu…
> I'm not surprised this isn't open to the public by Google yet, Closed models aren't going to matter in the long run. Hunyuan and LTX both run on consumer hardware and produce videos similar in quality to Sora Turbo, yet you can train them and prompt them on anything. They fit into the open source ecosystem which makes building plugins and controls super easy. Video is going to play out in a way that resembles image…
With the YouTube corpus at their disposal, I don't see how anyone can beat Google for AI video generation.
Re: Veo 2: Our video generation model
#123Namely, so few neurons to get picture in our heads.
I guess, end of the world scenarios may lead us to create that super intelligence with a gigantic ultra performant artificial "brain".
Re: Veo 2: Our video generation model
#124My theory as to why all the bigtech companies are investing so much money in video generation models is simple: they are trying to eliminate the threat of influencers/content creators to their ad revenue. Think about it, almost everyone I know rarely clicks on ads or buys from ads anymore. On the other hand, a lot of people including myself look into buying something advertised implicitly or explicitly by content cre…
Re: Veo 2: Our video generation model
#125Winning 2:1 in user preference versus sora turbo is impressive. It seems to have very similar limitations to sora. For example- the leg swapping in the ice skating video and the bee keeper picking up the jar is at a very unnatural acceleration (like it pops up). Though by my eye maybe slightly better emulating natural movement and physics in comparison to sora. The blog post has slightly more info: >at resolutions up…
Anyways, I strongly suspect that the funny meme content that seems to be the practical uses case of these video generators won't be possible on either Veo or Sora, because of copyright, PC, containing famous people, or other 'safety' related reasons.
Re: Veo 2: Our video generation model
#126We should collectively ignore these announcements of unavailable models. There are models you can use today, even in the EU.
Actually there is a pretty significant new model announced today and available now: "MiniMax (Hailuo)Video-01-Live" https://blog.fal.ai/introducing-minimax-hailuo-video-01-live... Although I tried that and it has the same issue all of them seem to have for me: if you are familiar with the face but they are not really famous then the features in the video are never close enough to be able to recognize the same person.
50 cents per video. Far more when accounting for a cherrypick rate.
Re: Veo 2: Our video generation model
#127Last time Google made a big Gemini announcement, OpenAI owned them by dropping the Sora preview shortly after. This feels like a bit of a comeback as Veo 2 (subjectively) appears to be a step up from what Sora is currently able to achieve.
Re: Veo 2: Our video generation model
#128This might be a dumb question to ask, but what exactly is this useful for? B-Roll for YouTube videos? I'm not sure why so much effort is being put into something like this when the applications are so limited.
If you want to train a model to have a general understanding of the physical world, one way is to show it videos and ask it to predict what comes next, and then evaluate it on how close it was to what actually came next. To really do well on this task, the model basically has to understand physics, and human anatomy, and all sorts of cultural things. So you're forcing the model to learn all these things about the wor…
All these models have just “seen” enough videos of all those things to build a probability distribution to predict the next step.
This is not bad, or make it inherently dumb, a major component of human intelligence is built on similar strategies. I couldn’t tell what grammatical rules are broken in text or what physical rules in a photograph but can tell it is wrong using the same methods .
Inference can take it far with large enough data sets, but sooner or later without reasoning you will hit a ceiling .
This is true for humans as well, plenty of people go far in life with just memorization and replication do a lot of jobs fairly competently, but not in everything.
Reasoning is essential for higher order functions and transformers is not the path for that
Re: Veo 2: Our video generation model
#129Earlier quoted context omitted.
Does everyone have "legal" access to YouTube. In theory that should matter to something like Open(Closed)Ai. But who knows.
I mean, I have trained myself on Youtube. Why can't a silicon being train itself on Youtube as well?