Live data from Hacker News

Sora: Creating video from text

openai.com

811–820 of 1001 posts

Re: Sora: Creating video from text

#811
post #639

I think the implications go much further than just the image/video considerations. This model shows a very good (albeit not perfect) understanding of the physics of objects and relationships between them. The announcement mentions this several times. The OpenAI blog post lists "Archeologists discover a generic plastic chair in the desert, excavating and dusting it with great care." as one of the "failed" cases. But t…

It doesn't understand physics. It just computes next frame based on current one and what it learned before, it's a plausible continuation. In the same way, ChatGPT struggles with math without code interpreter, Sora won't have accurate physics without a physics engine and rendering 3d objects. Now it's just a "what is the next frame of this 2D image" model plus some textual context.

[deleted]

Re: Sora: Creating video from text

#813
post #623

I think the implications go much further than just the image/video considerations. This model shows a very good (albeit not perfect) understanding of the physics of objects and relationships between them. The announcement mentions this several times. The OpenAI blog post lists "Archeologists discover a generic plastic chair in the desert, excavating and dusting it with great care." as one of the "failed" cases. But t…

I found the one about the people in Lagos pretty funny. The camera does about a 360deg spin in total, in the beginning there are markets, then suddenly there are skyscrapers in the background. So there's only very limited object permanence. > A beautiful homemade video showing the people of Lagos, Nigeria in the year 2056. Shot with a mobile phone camera. > https://cdn.openai.com/sora/videos/lagos.mp4

Also the women in red next to the people is very tiny and the market stall is also a mini market stall, and the table is made out of a bike.

For everyone that's carrying on about this thing understanding physics and has a model of the world...it's an odd world.

Re: Sora: Creating video from text

#814
post #406

This is insane. But I'm impressed most of all by the quality of motion . I've quite simply never seen convincing computer-generated motion before . Just look at the way the wooly mammoths connect with the ground, and their lumbering mass feels real. Motion-capture works fine because that's real motion, but every time people try to animate humans and animals, even in big-budget CGI movies, it's always ultimately obvio…

I disagree, just look at the legs of the woman in the first video. First she seems to be limping, than the legs rotate. The mammoth are totally uncanny for me as its both running and walking at the same time. Don't get me wrong, it is impressive. But I think many people will be very uncomfortable with such motion very quickly. Same story as the fingers before.

Yeah. I think people nowadays are in a kind of AI-euphoria and they took every advancement in AI for more than what they really are. The realization of their limitations will set in once people have been working long enough on the stuff. The capacity of the newfangled AIs are impressive. But even more impressive are their mimicry capabilities.

Re: Sora: Creating video from text

#815

I think the implications go much further than just the image/video considerations. This model shows a very good (albeit not perfect) understanding of the physics of objects and relationships between them. The announcement mentions this several times. The OpenAI blog post lists "Archeologists discover a generic plastic chair in the desert, excavating and dusting it with great care." as one of the "failed" cases. But t…

Wouldn't having a good understanding of physics mean you know that a women doesn't slide down the road when she walks? Wouldn't it know that a woolly mammoth doesn't emit profuse amounts steam when walking on frozen snow? Wouldn't the model know that legs are solid objects in which other object cannot pass through?

Maybe I'm missing the big picture here, but the above and all the weird spatial errors, like miniaturization of people make me think you're wrong.

Clearly the model is an achievement and doing something interesting to produce these videos, and they are pretty cool, but understanding physics seems like quite a stretch?

I also don't really get the excitement about the girl on the train in Tokyo:

In the Tokyo one, the model is smart enough to figure out that on a train, the reflection would be of a passenger, and the passenger has Asian traits since this is Tokyo

I don't know a lot about how this model works personally, but I'm guessing in the training data the vast majority of people riding trains in Tokyo featured asian people in them, assuming this model works on statistics like all of the other models I've seen recently from Open AI, then why is it interesting the girl in the reflection was Asian? Did you not expect that?

Re: Sora: Creating video from text

#818

Has anyone else noticed the leg swap in Tokyo video at 0:14. I guess we are past uncanny, but I do wonder if these small artifacts will always be present in generated content. Also begs the question, if more and more children are introduced to media from young age and they are fed more and more with generated content, will they be able to feel "uncanniness" or become completely blunt to it. There's definitely interes…

These artefacts go down with more compute. In four years when they attack it again with 100x compute and better algorithms I think it'll be virtually flawless.

Re: Sora: Creating video from text

#820
post #406

This is insane. But I'm impressed most of all by the quality of motion . I've quite simply never seen convincing computer-generated motion before . Just look at the way the wooly mammoths connect with the ground, and their lumbering mass feels real. Motion-capture works fine because that's real motion, but every time people try to animate humans and animals, even in big-budget CGI movies, it's always ultimately obvio…

I disagree, just look at the legs of the woman in the first video. First she seems to be limping, than the legs rotate. The mammoth are totally uncanny for me as its both running and walking at the same time. Don't get me wrong, it is impressive. But I think many people will be very uncomfortable with such motion very quickly. Same story as the fingers before.

Same story as the fingers before.

This is weird to me considering how much better this is than the SOTA still images 2 years ago. Even though there's weirdo artefacts in several of their example videos (indeed including migrating fingers), that stuff will be super easy to clean up, just as it is now for stills. And it's not going to stop improving.

Post reply on HN