Live data from Hacker News

Sora is here

openai.com

301–310 of 1001 posts

Re: Sora is here

#301

Earlier quoted context omitted.

> What will save it is that, no matter how picky you are as a creator, your audience will never know what exactly was that you dreamed up, so any half-decent approximation will work. Part of the problem is the "half decent approximations" tend towards a clichéd average, the audience won't know that the cool cyberpunk cityscape you generated isn't exactly what you had in mind, but they will know that it looks like eve…

> I think the pursuit of fidelity has made the models less creative over time (...) their output is ever more homogenized and interchangable. Ironically, we're long past that point with human creators , at least when it comes to movies and games. Take sci-fi movies, compare modern ones to the ones from the tail end of the 20th century. Year by year, VFX gets more and more detailed (and expensive) - more and better li…

Video games (the much larger industry of the two, by revenue) seems to be closer to understanding this. AAA games dominate advertising and news cycles, but on any best-seller list AAA games are on par with indie and B games (I think they call them AA now?). For every successful $60M PBR-rendered Unreal 5 title there is an equally successful game with low-fidelity graphics but exceptional art direction, story or gameplay.

Western movie studios may discover the same thing soon, with the number of high-budget productions tanking lately.

Re: Sora is here

#302
post #254

A little worried how young children watching these videos may develop inaccurate impressions of physics in nature. For instance, that ladybug looks pretty natural, but there's a little glitch in there that an unwitting observer, who's never seen a ladybug move before, may mistake as being normal. And maybe it is! And maybe it isn't? The sailing ship - are those water movements correct? The sinking of the elephant int…

I grew up watching Looney Tunes interpretation of physics and turned out just fine.

"A body at rest remains at rest until it looks down and realizes it has stepped off of a cliff."

Re: Sora is here

#303
post #149

Earlier quoted context omitted.

(2020) https://arxiv.org/abs/2010.11929 : an image is worth 16x16 words transformers for image recognition at scale (2021) https://arxiv.org/abs/2103.13915 : An Image is Worth 16x16 Words, What is a Video Worth? (2024) https://arxiv.org/abs/2406.07550 : An Image is Worth 32 Tokens for Reconstruction and Generation

Those are indeed 3 papers.

Yes in a nutshell they explain that you can express a picture or a video with relatively few discrete information.

First paper is the most famous and prompted a lot of research to using text generation tools in the image generation domain : 256 "words" for an image, Second paper is 24 reference image per minutes of video, Third paper is a refinement of the first saying you only need 32 "tokens". I'll let you multiply the numbers.

In kind of the same way as a who's who game, where you can identify any human on earth with ~32bits of information.

The corollary being that contrary to what parent is telling there is no theoretical obstacle to obtaining a video from a textual description.

Re: Sora is here

#304
post #135

I feel like there is a sweet spot for AI generation of images and videos that I would describe as "charmingly bad", like the stuff we got from the old CLIP+VQGAN models. I feel like Sora has jumped past that into the valley of "unappealingly bad".

I think that's why humor and memes are such good targets for this type of stuff. If you look up videos like "luma memes compilation," it takes well-known memes and distorts them in uncanny, freaky, and bizarre ways. Yet the fact the original subject is a meme somehow bypasses the uncanny valley repulsion. We seem to accept that much more readily, for whatever reason.

Re: Sora is here

#305
post #186

Earlier quoted context omitted.

AI isn't trying to sell to you: a precise artist with real vision in your brain. It is selling to managers who want to shit out something in an evening that approximates anything, that writes ads that no one wants to see anyway, that produces surface level examples of how you can pay employees less because "their job is so easy"

Yes and the thing is, even for those tasks, it's incredibly difficult to achieve even the low bar that a typical advertising manager expects. Try it yourself for any real world task and you will see.

Counterpoint: our CEO spent 25 minutes shitting out a bunch of AI ads because he was frustrated with the pace of our advertising creative team. They hated the ads that he created, for the reasons you mention, but we tested them anyways and the best performing ones beat all of our "expert" team's best ads by a healthy margin (on all the metrics we care about, from CTR to IPM and downstream stuff like retention and RoAS).

Maybe we're in a honeymoon period where your average user hasn't gotten annoyed by all the slop out there and they will soon, but at least for now, there is real value here. Yes, out of 20 ads maybe only 2 outperform the manually created ones, but if I can create those 20 with a couple hundred bucks in GenAI credits and maybe an hour or two of video editing that process wipes the floor with the competition, which is several thousand dollars per ad, most of which are terrible and end up thrown away, too. With the way the platforms function now, ad creative is quickly becoming a volume-driven "throw it at the wall and see what sticks" game, and AI is great for that.

Re: Sora is here

#307
post #80

Every day that passes I grow fonder of Google's decision to delay or otherwise keep a lot of this under the wraps. The other day I was scrolling down on YouTube shorts and a couple videos invoked an uncanny valley response from me (I think it was a clip of an unrealistically large snake covering some hut) which was somehow fascinating and strange and captivating, and then scrolling down a few more, again I saw someth…

I don't think Google delayed or kept this under wraps for any noble reasons. I think they were just disorganized as evidenced by their recent scrambling to compete in this space.

Re: Sora is here

#308

Earlier quoted context omitted.

> They may not match exactly what the writer had in mind, but they are fit for purpose. That's what GenAI is doing, too. After all, the audience only sees the final product; they never get know what the writer had in mind.

I haven't used SORA, but none of the GenAI I'm aware of could produce a competent comic book. When a human artist draws a character in a house in panel 1, they'll draw the same house in panel 2, not a procedurally generated different house for each image. If a 60 year old grizzled detective is introduced in page 1, a human artist will draw the same grizzled detective in page 2, 3 and so on, not procedurally generate…

A human artist keeps state :). They keep it between drawing sessions, and more importantly, they keep very detailed state - their imagination or interpretation of what the thing (house, grizzled detective, etc.) is.

Most models people currently use don't keep state between invocations, and whatever interpretation they make from provided context (e.g. reference image, previous frame) is surface level and doesn't translate well to output. This is akin to giving each panel in a comic to a different artist, and also telling them to sketch it out by their gut, without any deep analysis of prior work. It's a big limitation, alright, but researchers and practitioners are actively working to overcome it.

(Same applies to LLMs, too.)

Re: Sora is here

#309
post #80

Every day that passes I grow fonder of Google's decision to delay or otherwise keep a lot of this under the wraps. The other day I was scrolling down on YouTube shorts and a couple videos invoked an uncanny valley response from me (I think it was a clip of an unrealistically large snake covering some hut) which was somehow fascinating and strange and captivating, and then scrolling down a few more, again I saw someth…

I saw my first AI video that completely fooled commenters: https://imgur.com/a/cbjVKMU This was not marked as AI-generated and commenters were in awe at this fuzzy train, missing the "AIGC" signs. I'm quite nervous for the future.

The face of the girl on the left at the start in the first second should have been a giveaway.

Re: Sora is here

#310
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

The adage "a picture is worth a thousand words" has the nice corollary "A thousand words isn't enough to be precise about an image". Now expand that to movies and games and you can get why this whole generative-AI bubble is going to pop.

Sure it's going to pop. But when is the important question.

Being too early about this and being wrong are the same.

Post reply on HN