Live data from Hacker News

Sora is here

openai.com

51–60 of 1001 posts

Re: Sora is here

#51
I wonder what it is about EU and UK law, in particular, that restricts its availability there. Their FAQs don't mention this.

If it's about training models on potentially personal information, the GDPR (EU and UK variants) kicks in, but then that hasn't restricted OpenAI's ability to deploy (Chat)GPT there. The same applies to broader copyright regulations around platforms needing to proactively prevent copyright violation, something GPT could also theoretically accomplish. Any (planned) EU-specific regulations don't apply to the UK, so I doubt it's those either.

The only thing that leaves, perhaps, is laws around the generation of deepfakes which both the UK and EU have laws about? But then why didn't that affect DALL-E? Anyone with a more detailed understanding of this space have any ideas?

Re: Sora is here

#52

MKBHD's review of the new Sora release: https://www.youtube.com/watch?v=OY2x0TyKzIQ

Interesting to see how bad the physics/object permanence is. I wonder if combining this with a Genie 2 type model (Google's new "world model") would be the next step in refining it's capabilities.

Re: Sora is here

#53
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

Way back in the days of GPT-2, there was an expectation that you'd need to cherry-pick atleast 10% of your output to get something usable/coherent. GPT-3 and ChatGPT greatly reduced the need to cherry-pick, for better or for worse.

All the generated video startups seem to generate videos with much lower than 10% usable output, without significant human-guided edits. Given the massive amount of compute needed to generate a video relative to hyperoptimized LLMs, the quality issue will handicap gen video for the foreseeable future.

Re: Sora is here

#55

Yawn, there are literally 10 different apps and wannabe startups that do video generation and AI videos have already flooded social media. This doesn't look any better than what is and has been already available to the masses. OpenAI announced this ages ago and never did give people access, now competitors have already captured the AI generated video for social media slop market. We have yet to see any kind of AI cre…

Don't just critique - link. What other video generation tools have you used and recommend?

Re: Sora is here

#56

Not impressive compare to the opensource video models out there, I anticipated some physics/VR capabilities, but it's basically just a marketing promotion to "stay in the game"...

I... can you explain, or point to some competitors...? To me this looks leagues ahead of everything else. But maybe I'm behind the game?

AFAIK based on HuggingFace trending[1], the competitors are:

- bytedance/animatediff-lightning: https://arxiv.org/pdf/2403.12706 (2.7M downloads in the past 30d, released in March)

- genmo/mochi-1-preview: https://github-production-user-asset-6210df.s3.amazonaws.com... (21k downloads, released in October)

- thudm/cogvideox-5b: https://huggingface.co/THUDM/CogVideoX-5b (128k downloads, released in August)

Is there a better place to go? I'm very much not plugged into this part of LLMs, partially because it's just so damn spooky...

EDIT: I now see the reply above referencing Hunyuan, which I didn't even know was its own model. Fair enough! I guess, like always, we'll just need to wait for release so people can run their own human-preference tests to definitively say which is better. Hunyuan does indeed seem good

Re: Sora is here

#57

Earlier quoted context omitted.

Agreed. It’s still much better than what I could do myself without it, though. (Talking about visual generative AI in general)

Yeah, but if I handed you a Maxfield Parrish it would be better than either of us can do — but not what I asked for. I find generative AI frustrating because I know what I want. To this point I have been trying but then ultimately sitting it out — waiting for the one that really works the way I want.

For me even if I know what I want, if I’m using gen AI I’m happy to compromise and get good enough (which again, is so much better than I could do otherwise).

If you want higher quality/precision, you’ll likely want to ask a professional, and I don’t expect that to change in the near future.

Re: Sora is here

#58
If you're looking for video for casual personal projects or fill-ins for vlog posts, or something to make your PowerPoint look neat, this seems like a rad tool. It has a looong way to go before it's taking anyone's movie VFX job.

Re: Sora is here

#59
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

Not too far in the future you will be able to drag and drop the position of the characters as well as the position of the camera, among other refiment tools.

Re: Sora is here

#60
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

And another thing that irks me: none of these video generators get motion right...

Especially anything involving fluid/smoke dynamics, or fast dynamic momements of humans and animals all suffer from the same weird motion artifacts. I can't describe it other than that the fluidity of the movements are completely off.

And as all genai video tools I've used are suffering from the same problem, I wonder if this is somehow inherent to the approach & somehow unsolvable with the current model architectures.

Post reply on HN