I think the implications go much further than just the image/video considerations. This model shows a very good (albeit not perfect) understanding of the physics of objects and relationships between them. The announcement mentions this several times. The OpenAI blog post lists "Archeologists discover a generic plastic chair in the desert, excavating and dusting it with great care." as one of the "failed" cases. But t…
It doesn't understand physics. It just computes next frame based on current one and what it learned before, it's a plausible continuation. In the same way, ChatGPT struggles with math without code interpreter, Sora won't have accurate physics without a physics engine and rendering 3d objects. Now it's just a "what is the next frame of this 2D image" model plus some textual context.
Sora: Creating video from text
811–820 of 1001 posts
Re: Sora: Creating video from text
#812Re: Sora: Creating video from text
#813I think the implications go much further than just the image/video considerations. This model shows a very good (albeit not perfect) understanding of the physics of objects and relationships between them. The announcement mentions this several times. The OpenAI blog post lists "Archeologists discover a generic plastic chair in the desert, excavating and dusting it with great care." as one of the "failed" cases. But t…
I found the one about the people in Lagos pretty funny. The camera does about a 360deg spin in total, in the beginning there are markets, then suddenly there are skyscrapers in the background. So there's only very limited object permanence. > A beautiful homemade video showing the people of Lagos, Nigeria in the year 2056. Shot with a mobile phone camera. > https://cdn.openai.com/sora/videos/lagos.mp4
For everyone that's carrying on about this thing understanding physics and has a model of the world...it's an odd world.
Re: Sora: Creating video from text
#814This is insane. But I'm impressed most of all by the quality of motion . I've quite simply never seen convincing computer-generated motion before . Just look at the way the wooly mammoths connect with the ground, and their lumbering mass feels real. Motion-capture works fine because that's real motion, but every time people try to animate humans and animals, even in big-budget CGI movies, it's always ultimately obvio…
I disagree, just look at the legs of the woman in the first video. First she seems to be limping, than the legs rotate. The mammoth are totally uncanny for me as its both running and walking at the same time. Don't get me wrong, it is impressive. But I think many people will be very uncomfortable with such motion very quickly. Same story as the fingers before.
Re: Sora: Creating video from text
#815I think the implications go much further than just the image/video considerations. This model shows a very good (albeit not perfect) understanding of the physics of objects and relationships between them. The announcement mentions this several times. The OpenAI blog post lists "Archeologists discover a generic plastic chair in the desert, excavating and dusting it with great care." as one of the "failed" cases. But t…
Maybe I'm missing the big picture here, but the above and all the weird spatial errors, like miniaturization of people make me think you're wrong.
Clearly the model is an achievement and doing something interesting to produce these videos, and they are pretty cool, but understanding physics seems like quite a stretch?
I also don't really get the excitement about the girl on the train in Tokyo:
In the Tokyo one, the model is smart enough to figure out that on a train, the reflection would be of a passenger, and the passenger has Asian traits since this is Tokyo
I don't know a lot about how this model works personally, but I'm guessing in the training data the vast majority of people riding trains in Tokyo featured asian people in them, assuming this model works on statistics like all of the other models I've seen recently from Open AI, then why is it interesting the girl in the reflection was Asian? Did you not expect that?
Re: Sora: Creating video from text
#816Re: Sora: Creating video from text
#817Re: Sora: Creating video from text
#818Has anyone else noticed the leg swap in Tokyo video at 0:14. I guess we are past uncanny, but I do wonder if these small artifacts will always be present in generated content. Also begs the question, if more and more children are introduced to media from young age and they are fed more and more with generated content, will they be able to feel "uncanniness" or become completely blunt to it. There's definitely interes…
Re: Sora: Creating video from text
#819Re: Sora: Creating video from text
#820This is insane. But I'm impressed most of all by the quality of motion . I've quite simply never seen convincing computer-generated motion before . Just look at the way the wooly mammoths connect with the ground, and their lumbering mass feels real. Motion-capture works fine because that's real motion, but every time people try to animate humans and animals, even in big-budget CGI movies, it's always ultimately obvio…
I disagree, just look at the legs of the woman in the first video. First she seems to be limping, than the legs rotate. The mammoth are totally uncanny for me as its both running and walking at the same time. Don't get me wrong, it is impressive. But I think many people will be very uncomfortable with such motion very quickly. Same story as the fingers before.
This is weird to me considering how much better this is than the SOTA still images 2 years ago. Even though there's weirdo artefacts in several of their example videos (indeed including migrating fingers), that stuff will be super easy to clean up, just as it is now for stills. And it's not going to stop improving.