Live data from Hacker News

Sora is here

openai.com

231–240 of 1001 posts

Re: Sora is here

#231

Earlier quoted context omitted.

Do artist really have a fully formed vision in their head? I suspect the creative process is much more iterative rather than one-directional.

No one can have a fully formed vision. But intent, yes. Then you use techniques to materialize it. Word is a poor substitute for that intent, which is why there’s so many sketches in a visual project.

And why physical execution frequently significantly departs from sketches and concept art. The amount of intent that doesn't get translated is pretty staggering in both physical and digital pipelines in many projects.

Re: Sora is here

#232
post #124

Earlier quoted context omitted.

It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…

> Long term, you'll never have a coherent movie produced by stringing together a series of textual snippets because, again, that's just impossible. Why snippets? Submit a whole script the way a writer delivers a movie to a director. The (automated) director/DP/editor could maintain internal visual coherence, while the script drives the story coherence.

That's what I describe at the end, albeit quickly in lingo, where the internal coherence is maintained in internal embeddings that are never related to English at all. A top-level AI could orchestrate component AIs through embedded vectors, but you'll never do it with a human trying to type out descriptions.

Re: Sora is here

#233
post #164

I got lucky and got in moments after it launched, managed to get a video of "A pelican riding a bicycle along a coastal path overlooking a harbor" and then the queue times jumped up (my second video has been in the queue for 20+ minutes already) and the https://sora.com site now says "account creation currently unavailable" Here's my pelican video: https://simonwillison.net/2024/Dec/9/sora/

"The Pelican inexplicably morphs to cycle in the opposite direction half way through"

Oof, if sora can't even manage to maintain an internal consistency of the world for a 5 second short, I can't imagine how exacerbated it'll be at longer video generation times.

Re: Sora is here

#234
post #162

Anyone else find this stuff extremely distasteful? "Disrupting" creativity and art feels like it goes against our humanity.

"And then everyone clapped ..."

There's nothing wrong with technology going forward and this doesn't go against "creativity and art", to the contrary, it will enhance it.

Re: Sora is here

#235
A little worried how young children watching these videos may develop inaccurate impressions of physics in nature.

For instance, that ladybug looks pretty natural, but there's a little glitch in there that an unwitting observer, who's never seen a ladybug move before, may mistake as being normal. And maybe it is! And maybe it isn't?

The sailing ship - are those water movements correct?

The sinking of the elephant into snow - how deep is too deep? Should there be snow on the elephant or would it have melted from body heat? Should some of the snow fall off during movement or is it maybe packed down too tightly already?

There's no way to know because they aren't actual recordings, and if you don't know that, and this tech improves leaps and bounds (as we know it will), it will eventually become published and will be taken at face value by many.

Hopefully I'm just overthinking it.

Re: Sora is here

#236
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

I believe it. I was just using AI to help out with some mandatory end of year writing exercises at work.

Eventually, it starts to muck with the earlier work that it did good on, when I'm just asking it to add onto it.

I was still happy with what I got in the end, but it took trial and error and then a lot of piecemeal coaxing with verification that it didn't do more than I asked along the way.

I can imagine the same for video or images. You have to examine each step post prompt to verify it didn't go back and muck with the already good parts.

Re: Sora is here

#237

Earlier quoted context omitted.

The adage "a picture is worth a thousand words" has the nice corollary "A thousand words isn't enough to be precise about an image". Now expand that to movies and games and you can get why this whole generative-AI bubble is going to pop.

> Now expand that to movies and games and you can get why this whole generative-AI bubble is going to pop. What will save it is that, no matter how picky you are as a creator, your audience will never know what exactly was that you dreamed up, so any half-decent approximation will work. In other words, a corollary to your corollary is, "Fortunately, you don't need them to be, because no one cares about low-order bits…

That's just sad, and why people have a derogative stance towards generative AI: "half-decent" approximation removes all personality from the output, leading to a bunch of slop on the internet.

Re: Sora is here

#238
post #173

Earlier quoted context omitted.

something like a white paper with a mood board, color scheme, and concept art as the input might work. This could be sent into an LLM "expander" that increases the words and speficity. Then multiple reviews to tap things in the right direction.

And I think this realistically is going to be the shape of the tools to come in the foreseeable future.

You should see what people are building with Open Source video models like HunYuan [1] and ComfyUI + Control Nets. It blows Sora out of the water.

Check out the Banodoco Discord community [2]. These are the people pioneering steerable AI video, and it's all being built on top of open source.

[1] https://github.com/Tencent/HunyuanVideo

[2] https://banodoco.ai/

Re: Sora is here

#239
post #156

Earlier quoted context omitted.

Minimax (from China) and Kling 1.5 from China. Recently Tencent launched its own. You can see more model samples heee https://youtu.be/bCAV_9O1ioc

Those look... far worse? What am I missing.

Exactly I don't know how people are saying SORA is bad. I know there are restrictions with humans. But with the storyboard and other customisations, it's definitely up there!

Re: Sora is here

#240
post #199

Earlier quoted context omitted.

The past few years' innovation in AI has roughly been split into two camps for me. LLMs -- Awesome and useful. Disruptive, and somewhat dangerous, but probably more good than harm if we do it right. 'Generative art' (i.e. music generation, image generation, video generation) -- Why? Just why? The 'art' is always good enough to trick most humans at a glance but clearly fake, plastic, and soulless when you look a bit c…

There's this FinTech ad on the NYC subway right now. I can't remember the company, but the entire ad is just a picture of a guitar and some text. Anyway, the guitar is AI generated, and it's really bad. There are 5 strings, which morph into 6 at the headstock. There's a trem bar jammed under the pickguard, somehow. There's a randomly placed blob on the guitar that is supposed to be a knob/button, but clearly is not.…

Just need to add a hand with 6 fingers strumming it and it could be a meme.
Post reply on HN