This might be a dumb question to ask, but what exactly is this useful for? B-Roll for YouTube videos? I'm not sure why so much effort is being put into something like this when the applications are so limited.
Veo 2: Our video generation model
151–160 of 342 posts
Re: Veo 2: Our video generation model
#152This quote suggests not: "maintaining complete consistency throughout complex scenes or those with complex motion, remains a challenge."
Re: Veo 2: Our video generation model
#153Earlier quoted context omitted.
It doesn’t have to understand anything, none of these demonstrate reasoning or understanding. All these models have just “seen” enough videos of all those things to build a probability distribution to predict the next step. This is not bad, or make it inherently dumb, a major component of human intelligence is built on similar strategies. I couldn’t tell what grammatical rules are broken in text or what physical rule…
That's like saying that your brain doesn't understand anything, it just analyzes the visual data coming in via your eyes and predicts the next step of reality
we do extensive amount of pattern matching and drop enormous amount of sensory input very quickly because we expect patterns and assume a lot about our surroundings.
Unlearning this is a hard skill to pick up. There are many versions of training from martial arts to meditation that attempt to achieve this .
Point is that alone is not sufficient, the other core component is reasoning and understanding , transformers and learning on data is insufficient .
Parrot and few other animals can imitate human speech very well , that doesn’t mean they are understanding the speech or constructing .
Don’t get me wrong, i am not saying it is not useful, it is , but this attribution of reasoning and understanding to models that foundationally has no such building block is just being impressed by a speaking parrot
Re: Veo 2: Our video generation model
#154Earlier quoted context omitted.
Back when computers took up a whole room, you'd also have asked: "but what exactly is this useful for? B-Roll some simple calculations that anybody can do with a piece of paper and a pen."? Think 5-10 years into the future, this is a stepping stone
That's comparing apples to oranges though isn't it? Generating videos is the output of the technology, not the tech itself. It would be like someone asking "this computer that takes up a whole room printed out ascii art, what is this useful for?"
Re: Veo 2: Our video generation model
#155OpenAI is like the super luxurious yacht all pretty and shiny, while Google's AI department is the humongous nuclear submarine at least 5 times bigger than the yacht with a relatively cool conning tower, but not that spectacular to look at. Like the tanker which is still steering to fully align with the course people expect it to be, which they don't recognize that it will soon be there and be capable of rolling over…
Or, using Occams Razor; Sundar is a shit CEO and is playing catchup with a company largely fueled by innovations created at Google but never brought to market because it would eat into ads revenue. That, or they have a secret super human intelligence under wraps at the pentagon.
OpenAI might be well-capitalized, but they're (1) bleeding money, (2) no clear path to profitability, and (3) competing head-to-head with a behemoth who can profitably provide a similar offering at 10-20x cheaper (literally).
Google might be slow out the blocks, but it's not like they've been sitting on their hands for the past decade.
Re: Veo 2: Our video generation model
#156just to remind everyone that state of the art was Will Smith Eating Spaghetti in April of 2023 https://arstechnica.com/information-technology/2023/03/yes-v... We're not even done with 2024. Just imagine what's waiting for us in 2025.
Re: Veo 2: Our video generation model
#157Superficially impressive but what is the actual use case of the present state of the art? It makes 10-second demos, fine. But can a producer get a second shot of the same scene and the same characters, with visual continuity? Or a third, etc? In other words, can it be used to create a coherent movie --even a 60-second commercial -- with multiple shots having continuity of faces, backgrounds, and lighting? This quote…
Re: Veo 2: Our video generation model
#158Impressive but the page crashed chrome on my iPad!
Might be time for a new iPad. My old-school iPad Air has 2gb of memory and is an absolute hog when loading content-heavy websites.
The website is horrible on resources.
Re: Veo 2: Our video generation model
#159Website keeps crashing and reloading on Brave iOS.
Re: Veo 2: Our video generation model
#160Earlier quoted context omitted.
I mean, I have trained myself on Youtube. Why can't a silicon being train itself on Youtube as well?
Because silicon is a robot. A camcorder can't catch a flick with me in the theater even if I dress it up like a muppet.
A corporation "is a person" with all the rights that come along with that - free speech etc.