Live data from Hacker News

Sora is here

openai.com

401–410 of 1001 posts

Re: Sora is here

#401
“The version of Sora we are deploying has many limitations. It often generates unrealistic physics and struggles with complex actions over long durations. Although Sora Turbo is much faster than the February preview, we’re still working to make the technology affordable for everyone.”

So they demo the full model and release the quantised and censored model.

Does anyone else find this kind of bait & switch distasteful?

Re: Sora is here

#402

Earlier quoted context omitted.

I saw my first AI video that completely fooled commenters: https://imgur.com/a/cbjVKMU This was not marked as AI-generated and commenters were in awe at this fuzzy train, missing the "AIGC" signs. I'm quite nervous for the future.

The face of the girl on the left at the start in the first second should have been a giveaway.

Hard to not discount that as a compression artifact.

Re: Sora is here

#403
post #300
post #212

Earlier quoted context omitted.

It saddens me. Innovations in AI 'art' generation (music, audio, photo) have been a net negative to society and are already actively harming the Internet and our media sphere. Like I said in another comment, LLMs are cool and useful, but who in the hell asked for AI art? It's good enough to fool people and break the fragile trust relationship we had with online content, but is also extremely shit and carries no meani…

I think AI "art" can be as useful as the text generators, i.e. only within certain limits of dull and stupid stuff that needs to exist but has little to no value. For example, you need to generate a landing page for your boring company: text, images, videos and the overall design (as well as code!) can be and should be generated because... who cares about your boring company's landing page, right?

One could ask why the boring company landing page exists in the first place though. If it's not providing value to humans to warrant actual attention being paid to it...

Re: Sora is here

#404

Earlier quoted context omitted.

VFX heavy feature for a Disney subsidiary. Each frame is rendered independently of each other - it’s not like video encoding where each frame depends on the previous one, they all have their own scene assembly that can be sent to a server to parallelize rendering. With enough compute, the entire film can be rendered in a few days. (It’s a little more complicated than that but works to a first order approximation) I d…

Are they still using CPUs and not GPUs for rendering? Weren't the rendering algos ported to CUDA yet?

GPU renderers exist but they have pretty hard scaling limits, so the highest end productions still use CPU renderers almost exclusively.

The 3D you see in things like commercials is usually done on GPUs though because at their smaller scale it's much faster.

Re: Sora is here

#405
post #124

Earlier quoted context omitted.

It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…

> Long term, you'll never have a coherent movie produced by stringing together a series of textual snippets because, again, that's just impossible. Why snippets? Submit a whole script the way a writer delivers a movie to a director. The (automated) director/DP/editor could maintain internal visual coherence, while the script drives the story coherence.

You should watch how movies are made sometime. How a script is developed. How changes to it are made. How storyboards are created. How actors are screened for roles. How locations are scouted, booked, and changed. How the gazillion of different departments end up affecting how a movie looks, is produced, made, and in which direction it goes (the wardrobe alone, and its availability and deadlines will have a huge impact on the movie).

What does "EXT. NIGHT" mean in a script? Is it cloudy? Rainy? Well lit? What are camera locations? Is the scene important for the context of the movie? What are characters wearing? What are they looking at?

What do actors actually do? How do they actually behave?

Here are a few examples of script vs. screen.

Here's a well described script of Whiplash. Tell me the one hundred million things happening on screen that are not in the script: https://www.youtube.com/watch?v=kunUvYIJtHM

Or here's Joker interrogation from The Dark Night Rises. Same million different things, including actors (or the director) ignoring instructions in the script: https://www.youtube.com/watch?v=rqQdEh0hUsc

Here's A Few Good Men: https://www.youtube.com/watch?v=6hv7U7XhDdI&list=PLxtbRuSKCC...

and so on

---

Edit. Here's Annie Atkins on visual design in movies, including Grand Budapest Hotel: https://www.youtube.com/watch?v=SzGvEYSzHf4. And here's a small article summarizing some of it: https://www.itsnicethat.com/articles/annie-atkins-grand-buda...

Good luck finding any of these details in any of the scripts. See minute 14:16 where she goes through the script

Edit 2: do watch The Kerning chapter at 22:35 to see what it actually takes to create something :)

Re: Sora is here

#406
post #360

Earlier quoted context omitted.

Yes in a nutshell they explain that you can express a picture or a video with relatively few discrete information. First paper is the most famous and prompted a lot of research to using text generation tools in the image generation domain : 256 "words" for an image, Second paper is 24 reference image per minutes of video, Third paper is a refinement of the first saying you only need 32 "tokens". I'll let you multiply…

I think something is getting lost in translation. These papers, from my quick skim (tho I did read the first one fully years ago,) seem to show that some images and to an extent video can be generated from discrete tokens, but does not show that exact images nor that any image can be. For instance, what combination of tokens must I put in to get _exactly_ Mona Lisa or starry night? (Tho these might be very well repre…

If you want to know what tokens you want to obtain _exactly_ Mona Lisa, or any other image, you take the image and put it through your image tokenizer aka encode it, and if you have the sequence of token you can decode it to an image.

VQ-VAE (Vector Quantised-Variational AutoEncoder), (2017) https://arxiv.org/abs/1711.00937

The whole encoding-decoding process is reversible, and you only lose some imperceptible "details", the process can be either trained with a L2Loss, or a perceptual loss depending what you value.

The point being that images which occurs naturally are not really information rich and can be compressed a lot by neural networks of a few GB that have seen billions of pictures. With that strong prior, aka common knowledge, we can indeed paint with words.

Re: Sora is here

#408

Earlier quoted context omitted.

The adage "a picture is worth a thousand words" has the nice corollary "A thousand words isn't enough to be precise about an image". Now expand that to movies and games and you can get why this whole generative-AI bubble is going to pop.

If you can build a system that can generate engaging games and movies, from an economic (bubble popping or not popping) point of view it's largely irrelevant whether they conform to fine-grained specifications by a human or not.

Text generation is the most mature form of genAI and even that isn't even remotely close to producing infinite engaging stories. Adding the visual aspect to make that story into a movie or the interactive element to turn it into a game is only uphill from there.

Re: Sora is here

#409
post #80

Every day that passes I grow fonder of Google's decision to delay or otherwise keep a lot of this under the wraps. The other day I was scrolling down on YouTube shorts and a couple videos invoked an uncanny valley response from me (I think it was a clip of an unrealistically large snake covering some hut) which was somehow fascinating and strange and captivating, and then scrolling down a few more, again I saw someth…

Considering google image search is polluted by AI-generated images at this moment, perhaps google is afraid of making the search even worse?

Re: Sora is here

#410
post #124
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…

Can't you just give it a photo of a dog, and then say "use this dog in this or that scene"?
Post reply on HN