Live data from Hacker News

Sora is here

openai.com

211–220 of 1001 posts

Re: Sora is here

#211
post #164

I got lucky and got in moments after it launched, managed to get a video of "A pelican riding a bicycle along a coastal path overlooking a harbor" and then the queue times jumped up (my second video has been in the queue for 20+ minutes already) and the https://sora.com site now says "account creation currently unavailable" Here's my pelican video: https://simonwillison.net/2024/Dec/9/sora/

For those who can't try Sora out, Tencent's super recent HunYuan is 100% open source and outperforms Sora. It's compatible with fine tuning, ComfyUI development, and is getting all manner of ControlNets and plugins.

I don't see how Sora can stay in this race. The open source commoditization is going to hit hard, and OpenAI probably doesn't have the product DNA or focus to bark up this tree too.

Tencent isn't the only company releasing open weights. Genmo, Black Forest Labs, and Lightricks are developing completely open source video models, and that's .

Even if there weren't open source competitors, there are a dozen closed source foundation video companies: Runway, Pika, Kling, Hailuo, etc.

I don't think OpenAI can afford to divert attention and win in this space. It'll be another Dall-E vs. Midjourney, Flux, Stable Diffusion.

https://github.com/Tencent/HunyuanVideo

https://x.com/kennethlynne/status/1865528133807386666

https://fal.ai/models/fal-ai/hunyuan-video

Re: Sora is here

#212
post #80

Every day that passes I grow fonder of Google's decision to delay or otherwise keep a lot of this under the wraps. The other day I was scrolling down on YouTube shorts and a couple videos invoked an uncanny valley response from me (I think it was a clip of an unrealistically large snake covering some hut) which was somehow fascinating and strange and captivating, and then scrolling down a few more, again I saw someth…

It saddens me. Innovations in AI 'art' generation (music, audio, photo) have been a net negative to society and are already actively harming the Internet and our media sphere.

Like I said in another comment, LLMs are cool and useful, but who in the hell asked for AI art? It's good enough to fool people and break the fragile trust relationship we had with online content, but is also extremely shit and carries no meaning or depth whatsoever.

Re: Sora is here

#213
post #164

I got lucky and got in moments after it launched, managed to get a video of "A pelican riding a bicycle along a coastal path overlooking a harbor" and then the queue times jumped up (my second video has been in the queue for 20+ minutes already) and the https://sora.com site now says "account creation currently unavailable" Here's my pelican video: https://simonwillison.net/2024/Dec/9/sora/

Thanks, would you mind elaborate more on what you wrote below:

  Sora is built entirely around the idea of directly manipulating and editing and remixing the clips it generates, so the goal isn't to have it produce usable videos from a single prompt.

Re: Sora is here

#214
post #164

I got lucky and got in moments after it launched, managed to get a video of "A pelican riding a bicycle along a coastal path overlooking a harbor" and then the queue times jumped up (my second video has been in the queue for 20+ minutes already) and the https://sora.com site now says "account creation currently unavailable" Here's my pelican video: https://simonwillison.net/2024/Dec/9/sora/

> The Pelican inexplicably morphs to cycle in the opposite direction half way through

It's pretty cool though, the kind of thing that'd be hard if it was what you actually wanted!

Re: Sora is here

#215
post #124
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…

> Long term, you'll never have a coherent movie produced by stringing together a series of textual snippets because, again, that's just impossible.

Why snippets? Submit a whole script the way a writer delivers a movie to a director. The (automated) director/DP/editor could maintain internal visual coherence, while the script drives the story coherence.

Re: Sora is here

#216

Wow this is bad. And by bad i mean worse than leading open source and existing alternatives. Is it me or does it seem like OpenAI revolutionized with both chatGPT and Sora, but they've completely hit the ceiling? Honestly a bit surprised it happened so fast!

What are the leading alternatives? (Open source or otherwise)

You have to be specific. What's more important to you?

- uncensored output (SD + LoRa)

- Overall speed of generation (midjourney)

- Image quality (probably midjourney, or an SDXL checkpoint + upscaler)

- Prompt adherence (flux, DALL-E 3)

EDIT: This is strictly around image generation. The main video competitors are Kling, Hailuo, and Runway.

Re: Sora is here

#217
post #164

I got lucky and got in moments after it launched, managed to get a video of "A pelican riding a bicycle along a coastal path overlooking a harbor" and then the queue times jumped up (my second video has been in the queue for 20+ minutes already) and the https://sora.com site now says "account creation currently unavailable" Here's my pelican video: https://simonwillison.net/2024/Dec/9/sora/

That's an awful result. It turning around has absolutely nothing to do with what you asked for. It's similar in nature to what the chatbot in the recent and ongoing scandal said, saying to come home to her, when it should have known that the idea would be nonsensical or could be taken to mean something horrendous. https://apnews.com/article/chatbot-ai-lawsuit-suicide-teen-a...

So you were lucky indeed to be able to run your prompt and share it, because the result was quite illuminating, but not in a way that looks good for Sora and OpenAI as a whole.

Re: Sora is here

#218

Earlier quoted context omitted.

> Now expand that to movies and games and you can get why this whole generative-AI bubble is going to pop. What will save it is that, no matter how picky you are as a creator, your audience will never know what exactly was that you dreamed up, so any half-decent approximation will work. In other words, a corollary to your corollary is, "Fortunately, you don't need them to be, because no one cares about low-order bits…

Do artist really have a fully formed vision in their head? I suspect the creative process is much more iterative rather than one-directional.

No one can have a fully formed vision. But intent, yes. Then you use techniques to materialize it. Word is a poor substitute for that intent, which is why there’s so many sketches in a visual project.

Re: Sora is here

#219
post #169

Earlier quoted context omitted.

How does your conclusion follow from your statement? Neural networks are largely black box piles of linear algebra which are massaged to minimize a loss function. How would you incorporate smooth kinematic motion in such an environment? The fact that you discount the knowledge of literally every single employee at OpenAI is a big signal that you have no idea what you’re talking about. I don’t even really like OpenAI…

I've seen the quality of OpenAI engineers on Twitter and it's easy enough to extrapolate. Moreoever, neural networks are not black boxes, you're just parroting whatever you've heard on social media. The underlying theory is very simple.

Do not make assumptions about people you do not know in an attempt to discredit them. You seem to be a big fan of that.

I have been working with NLP and neural networks since 2017.

They aren’t just black boxes, they are _largely_ black boxes.

When training an NN, you don’t have great control over what parts of the model does what or how.

Now instead of trying to discredit me, would you mind answering my question? Especially since, as you say, the theory is so simple.

How would you incorporate smooth kinematic motion in such an environment?

Re: Sora is here

#220
post #164

I got lucky and got in moments after it launched, managed to get a video of "A pelican riding a bicycle along a coastal path overlooking a harbor" and then the queue times jumped up (my second video has been in the queue for 20+ minutes already) and the https://sora.com site now says "account creation currently unavailable" Here's my pelican video: https://simonwillison.net/2024/Dec/9/sora/

I don't have a lot of mental model for how this works, but I was surprised to note that it seems to maintain continuity on the shapes of the bushes and brown spots on the grass that track out of frame on the left and then reappear as it pans back into frame.
Post reply on HN