Live data from Hacker News

Sora: Creating video from text

openai.com

401–410 of 1001 posts

Re: Sora: Creating video from text

#402

This is all very impressive. I can't help to wonder though. How is text-to-video going to benefit humanity? That's what OpenAI is supposedly about, right? We'll get some groundbreaking film content out of this in the hands of a few talented creatives, and a vast ocean of mediocre content from the hands of talentless people who know how to type. What's the benefit to humanity, concretely?

If a model can generate it, it can understand it. They can probably reverse engineer this to build a multi-modal GPT that is fed video and understands what is going on. That's how you get "smart" robots. Active scene understanding via the video modality + conversational capabilities via the text/audio modality.

But we can already do this?

Re: Sora: Creating video from text

#404

This is all very impressive. I can't help to wonder though. How is text-to-video going to benefit humanity? That's what OpenAI is supposedly about, right? We'll get some groundbreaking film content out of this in the hands of a few talented creatives, and a vast ocean of mediocre content from the hands of talentless people who know how to type. What's the benefit to humanity, concretely?

> Sora serves as a foundation for models that can understand and simulate the real world, a capability we believe will be an important milestone for achieving AGI.

Re: Sora: Creating video from text

#405
Not that this isn't a leaps and bounds improvement over the state of the art, but it's interesting to look at the mistakes it makes - where do we still need improvements?

This video is pretty instructive: https://cdn.openai.com/sora/videos/amalfi-coast.mp4

It "eats" several people with the wall part of the way through the video, and the camera movements are odd. Strange camera movements, in response to most of the prompts, seems like the biggest problem. The model arbitrarily decides to change direction on a dime - even a drone wouldn't behave quite like that.

Re: Sora: Creating video from text

#406

This is insane. But I'm impressed most of all by the quality of motion . I've quite simply never seen convincing computer-generated motion before . Just look at the way the wooly mammoths connect with the ground, and their lumbering mass feels real. Motion-capture works fine because that's real motion, but every time people try to animate humans and animals, even in big-budget CGI movies, it's always ultimately obvio…

I disagree, just look at the legs of the woman in the first video. First she seems to be limping, than the legs rotate. The mammoth are totally uncanny for me as its both running and walking at the same time.

Don't get me wrong, it is impressive. But I think many people will be very uncomfortable with such motion very quickly. Same story as the fingers before.

Re: Sora: Creating video from text

#407
post #381

This is insane. But I'm impressed most of all by the quality of motion . I've quite simply never seen convincing computer-generated motion before . Just look at the way the wooly mammoths connect with the ground, and their lumbering mass feels real. Motion-capture works fine because that's real motion, but every time people try to animate humans and animals, even in big-budget CGI movies, it's always ultimately obvio…

When others create text to video systems (eg. Lumiere from Google) they publish the research (eg. https://arxiv.org/pdf/2401.12945.pdf ). Open AI is all about commercialization. I don't like their attitude

OAI requires a real mobile phone number to signup and are therefore an adtech company.

Re: Sora: Creating video from text

#408
post #381

This is insane. But I'm impressed most of all by the quality of motion . I've quite simply never seen convincing computer-generated motion before . Just look at the way the wooly mammoths connect with the ground, and their lumbering mass feels real. Motion-capture works fine because that's real motion, but every time people try to animate humans and animals, even in big-budget CGI movies, it's always ultimately obvio…

When others create text to video systems (eg. Lumiere from Google) they publish the research (eg. https://arxiv.org/pdf/2401.12945.pdf ). Open AI is all about commercialization. I don't like their attitude

Not to be overly cute, but if the cutting edge research you do is maybe changing the world fundamentally, forever, guarding that tech should be really, really, really far up your list of priorities and everyone else should be really happy about your priorities.

And that should probably take precedence over the semantics of your moniker, every single time (even if hn continues to be super sour about it)

Re: Sora: Creating video from text

#409
Does anyone know how to handle the depression/doom one feels with these updates?

Yes, it's a great technical achievement, but I just worry for the future. We don't have good social safety nets, and we aren't close to UBI. It's difficult for me to see that happen unless something drastic changes.

I'm also afraid of one company just having so much power. How does anyone compete?

Re: Sora: Creating video from text

#410

I wonder what served as the dataset for the model. Videos on YouTube presumably, since messing around with the film industry would be too expensive?

Almost certainly troves of stock footage. The type of exaggerated motion seen in these examples is very reminiscent of stock footage. And it is heavily textually annotated for search.
Post reply on HN