In less than a few hours Gemini 1.5 is old news. Sam is doing live demos on Twitter while Google just released a blog. Didn't think Google would be the first of the Facebook, Apple, Google and Microsoft to get disrupted.
Did they have this ready to go to upstage whatever Google would release? Or just coincidental both things announced today?
Sora: Creating video from text
361–370 of 1001 posts
Re: Sora: Creating video from text
#362Re: Sora: Creating video from text
#363This model shows a very good (albeit not perfect) understanding of the physics of objects and relationships between them. The announcement mentions this several times.
The OpenAI blog post lists "Archeologists discover a generic plastic chair in the desert, excavating and dusting it with great care." as one of the "failed" cases. But this (and "Reflections in the window of a train traveling through the Tokyo suburbs.") seem to me to be 2 of the most important examples.
- In the Tokyo one, the model is smart enough to figure out that on a train, the reflection would be of a passenger, and the passenger has Asian traits since this is Tokyo. - In the chair one, OpenAI says the model failed to model the physics of the object (which hints that it did try to, which is not how the early diffusion models worked ; they just tried to generate "plausible" images). And we can see one of the archeologists basically chasing the chair down to grab it, which does correctly model the interaction with a floating object.
I think we can't underestimate how crucial that is to the building of a general model that has a strong model of the world. Not just a "theory of mind", but a litteral understanding of "what will happen next", independently of "what would a human say would happen next" (which is what the usual text-based models seem to do).
This is going to be much more important, IMO, than the video aspect.
Re: Sora: Creating video from text
#364Many might miss the key paragraph at the end: "Sora serves as a foundation for models that can understand and simulate the real world, a capability we believe will be an important milestone for achieving AGI." This also helps explain why the model is so good since it is trained to simulate the real world, as opposed to imitate the pixels. More importantly, its capabilities suggest AGI and general robotics could be cl…
> since it is trained to simulate the real world Is it though? Or is this just marketing?
Re: Sora: Creating video from text
#365Obviously incredibly cool, but it seems that people are incredibly overstating the applications of this. Realistically, how do you fit this into a movie, a TV show, or a game? You write a text prompt, get a scene, and then everything is gone—the characters, props, rooms, buildings, environments, etc. won’t carry over to the next prompt.
Re: Sora: Creating video from text
#366In less than a few hours Gemini 1.5 is old news. Sam is doing live demos on Twitter while Google just released a blog. Didn't think Google would be the first of the Facebook, Apple, Google and Microsoft to get disrupted.
I mean, why would this make google look bad? Gemini is catching up, so OpenAI needs a new venue to market itself to the investors. It is doing a soft pivoting if you ask me, now GPT4 is like not that special anymore.
Re: Sora: Creating video from text
#367https://openai.com/sora?video=big-sur In this video, there's extremely consistent geometry as the camera moves, but the texture of the trees/shrubs on the top of the cliff on the left seems to remain very flat, reminiscent of low-poly geometry in games. I wonder if this is an artifact of the way videos are generated. Is the model separating scene geometry from camera? Maybe some sort of video-NeRF or Gaussian Splatti…
Edit: Here[0] I highlighted a groove in the bushes moving with perfect perspective
Re: Sora: Creating video from text
#368Re: Sora: Creating video from text
#369I do wonder why OpenAI chose the name "Sora" for this model. AI is now going to have intersectionality with Kingdom Hearts. (Atleast you don't need a PhD to understand AI.)
Re: Sora: Creating video from text
#370This is leaps and bounds beyond anything out there, including both public models like SVD 1.1 and Pika Labs' / Runway's models. Incredible.