Live data from Hacker News

Sora is here

openai.com

111–120 of 1001 posts

Re: Sora is here

#112
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

And another thing that irks me: none of these video generators get motion right... Especially anything involving fluid/smoke dynamics, or fast dynamic momements of humans and animals all suffer from the same weird motion artifacts. I can't describe it other than that the fluidity of the movements are completely off. And as all genai video tools I've used are suffering from the same problem, I wonder if this is someho…

I think one of the biggest problems is the models are trained on 2D sequences and don't have any understanding of what they're actually seeing. They see some structure of pixels shift in a frame and learn that some 2D structures should shift in a frame over time. They don't actually understand the images are 2D capture of an event that occurred in four dimensions and the thing that's been imaged is under the influence of unimaged forces.

I saw a Santa dancing video today and the suspension of disbelief was almost instantly dispelled when the cuffs of his jacket moved erratically. The GenAI was trying to get them to sway with arm movements but because it didn't understand why they would sway it just generated a statistical approximation of swaying.

GenAI also definitely doesn't understand 3D structures easily demonstrated by completely incorrect morphological features. Even my dogs understand gravity, if I drop an object they're tracking (food) they know it should hit the ground. They also understand 3D space, if they stand on their back legs they can see over things or get a better perspective.

I've yet to see any GenAI that demonstrates even my dogs' level of understanding the physical world. This leaves their output in the uncanny valley.

Re: Sora is here

#113
post #17

Not available in the EU: https://help.openai.com/en/articles/10250692-sora-supported-...

Does VPN solves the problem? I'm living in an EU country and I don't like that the EU decides for me (and companies like OpenAI or Meta don't give out their models to me)! I'm an old enough adult to decide for myself what I want...

Re: Sora is here

#114

Earlier quoted context omitted.

Yeah, but if I handed you a Maxfield Parrish it would be better than either of us can do — but not what I asked for. I find generative AI frustrating because I know what I want. To this point I have been trying but then ultimately sitting it out — waiting for the one that really works the way I want.

For me even if I know what I want, if I’m using gen AI I’m happy to compromise and get good enough (which again, is so much better than I could do otherwise). If you want higher quality/precision, you’ll likely want to ask a professional, and I don’t expect that to change in the near future.

That limits its value for industries like Hollywood, though, doesn't it? And without that, who exactly is going to pay for this?

Re: Sora is here

#115

Hollywood's days are numbered. If you are a creative in this industry, start preparing to transition to another industry or adapt. Your boss is highly likely to be toying around with this. The first entirely AI generated film (with Sora or other AI video tools) to win an Oscar will be less than 5 years away.

I'd take that bet at 10:1 odds.

I'd be careful.

OpenAI could be a big enough bubble in less than 5 years to buy the Oscar winner, even if the film is terrible.

Also, OP only said "an Oscar".

The Oscar committee could easily get themselves hyped enough on the AI bubble, to create an AI Oscar Film award.

No one said anything about making a "good" movie.

Re: Sora is here

#116
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

Real artists struggle matching vague descriptions of what is in your head too. This is at least quicker?

Real artists take comic book scripts and turn them into actual comic books every month. They may not match exactly what the writer had in mind, but they are fit for purpose.

Re: Sora is here

#117

Hollywood's days are numbered. If you are a creative in this industry, start preparing to transition to another industry or adapt. Your boss is highly likely to be toying around with this. The first entirely AI generated film (with Sora or other AI video tools) to win an Oscar will be less than 5 years away.

Nothing I'm seeing here looks like it's going to destroy Hollywood.

I could see this tool maybe being used for generating establishing shots (generate a sweeping drone shot of a lighthouse looking out over a stormy sea), but then the actual talent work in a scene will be way more sensitive. The little details matter so much, and this feels so far from getting all of that right.

Sure, this is the worst it will ever be, things will improve, etc, but if we've learned anything with AI, it's that the last mile is often the hardest.

Re: Sora is here

#118

Earlier quoted context omitted.

And another thing that irks me: none of these video generators get motion right... Especially anything involving fluid/smoke dynamics, or fast dynamic momements of humans and animals all suffer from the same weird motion artifacts. I can't describe it other than that the fluidity of the movements are completely off. And as all genai video tools I've used are suffering from the same problem, I wonder if this is someho…

Neural networks use smooth manifolds as their underlying inductive bias so in theory it should be possible to incorporate smooth kinematic and Hamiltonian constraints but I am certain no one at OpenAI actually understands enough of the theory to figure out how to do that.

> I am certain no one at OpenAI actually understands enough of the theory to figure out how to do that

We would love to learn more about the origin of your certainty.

Re: Sora is here

#119
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

Way back in the days of GPT-2, there was an expectation that you'd need to cherry-pick atleast 10% of your output to get something usable/coherent. GPT-3 and ChatGPT greatly reduced the need to cherry-pick, for better or for worse. All the generated video startups seem to generate videos with much lower than 10% usable output, without significant human-guided edits. Given the massive amount of compute needed to gener…

Plus editing text or an image is practical. Video editors typically are used to cut and paste video streams - a video editor can't fix a stream of video that gets motion or anatomy wrong.

Re: Sora is here

#120
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

The adage "a picture is worth a thousand words" has the nice corollary "A thousand words isn't enough to be precise about an image". Now expand that to movies and games and you can get why this whole generative-AI bubble is going to pop.

(2020) https://arxiv.org/abs/2010.11929 : an image is worth 16x16 words transformers for image recognition at scale

(2021) https://arxiv.org/abs/2103.13915 : An Image is Worth 16x16 Words, What is a Video Worth?

(2024) https://arxiv.org/abs/2406.07550 : An Image is Worth 32 Tokens for Reconstruction and Generation

Post reply on HN