Live data from Hacker News

Sora: Creating video from text

openai.com

801–810 of 1001 posts

Re: Sora: Creating video from text

#801

Does anyone else feel a sense of doom from these advancements? I'm definitely not a Luddite, I've been working professionally as a programmer for quite some time now, but I just can't shake this feeling. And this is not in the "I might lose my job to this" kind of feeling, that's obviously there, but it's something deeper, more sinister. I don't think I can explain it properly. Anyway, videos look incredible. I genui…

The compute and innovations behind it should be owned by the planet, not by a handful of billionaires. It is far too powerful to be controlled by such a small group of humans, who decide what is "safe" and what isn't. It took billions of years for all of our ancestors to enable this technology, and now a handful claim it for themselves. The GPUs to run these models cost $20,000+ each, and only the ultra-rich can affo…

Actually, you can live-render around 12fps videos on a consumer gaming rig using software installable in a night ($3k). Not as fancy internally-consistent videos as these, but still impressive - and that's just an algorithm update and model download away. And every second a corporate AI model is exposed publicly to the world that's more training that can be siphoned to open source models at far more cost effective rates than the initial leaders.

You're impressed by the lions. But us hyenas and vultures will get our turns still too. This is not over. Information innately diffuses.

We need to organize, and we need to build.

Re: Sora: Creating video from text

#803
post #639

I think the implications go much further than just the image/video considerations. This model shows a very good (albeit not perfect) understanding of the physics of objects and relationships between them. The announcement mentions this several times. The OpenAI blog post lists "Archeologists discover a generic plastic chair in the desert, excavating and dusting it with great care." as one of the "failed" cases. But t…

It doesn't understand physics. It just computes next frame based on current one and what it learned before, it's a plausible continuation. In the same way, ChatGPT struggles with math without code interpreter, Sora won't have accurate physics without a physics engine and rendering 3d objects. Now it's just a "what is the next frame of this 2D image" model plus some textual context.

GPT-4 doesn't "struggle with math". It does fine. Most humans aren't any better.

Sora is not autoregressive anyway but there's nothing "just" and next frame/token prediction.

Re: Sora: Creating video from text

#804

Does anyone else feel a sense of doom from these advancements? I'm definitely not a Luddite, I've been working professionally as a programmer for quite some time now, but I just can't shake this feeling. And this is not in the "I might lose my job to this" kind of feeling, that's obviously there, but it's something deeper, more sinister. I don't think I can explain it properly. Anyway, videos look incredible. I genui…

I think of it like: The only reason humans still drive cars is we have yet to find a good enough way of replacing ourselves with something more effective. It's merely an implementation detail of "getting from A to B" that would be disrupted if a true autonomous solution was discovered. Many would want to optimize away drunk drivers and road rage if it were possible in some faraway future. So something like a steering wheel could be seen like a compromise of sorts, until the next big thing makes them obsolete.

That, and the state of missing a technology in a period of time is irreplaceable once it's been discovered. Nobody can live in an era without social media anymore, barring a global-scale catastrophic reset. So I believe it's important to consider what technology is not yet totally pervasive, for example by realizing there is still a steering wheel for you to grip in your car.

And in my mind, the sinister feeling stems from the fact that all it takes to irreversibly shift society like that is enough smart people with honest intentions but little foresight of what will happen in a few decades as a result of proliferating all this. The problems that result stop being in anyone's control, "throwing it over the wall" so to speak, and instead become yet another fact of life that could weigh us down (mostly I think of the ubiquity of social media and how it has changed human interaction). And it all stems from just a few engineering type people getting overexcited about cool possibilities they can grasp at, not considering there are billions of people unlike them who may have other ideas.

Re: Sora: Creating video from text

#806

Does anyone else feel a sense of doom from these advancements? I'm definitely not a Luddite, I've been working professionally as a programmer for quite some time now, but I just can't shake this feeling. And this is not in the "I might lose my job to this" kind of feeling, that's obviously there, but it's something deeper, more sinister. I don't think I can explain it properly. Anyway, videos look incredible. I genui…

[deleted]

Re: Sora: Creating video from text

#807
post #49

https://openai.com/sora?video=big-sur In this video, there's extremely consistent geometry as the camera moves, but the texture of the trees/shrubs on the top of the cliff on the left seems to remain very flat, reminiscent of low-poly geometry in games. I wonder if this is an artifact of the way videos are generated. Is the model separating scene geometry from camera? Maybe some sort of video-NeRF or Gaussian Splatti…

It looks perfect to me. That's exactly how the area looks in person.

Re: Sora: Creating video from text

#808

Watching these made me think, I'm going to want to go to the theatre a lot more in the future and see fellow humans in plays, lectures and concerts. Such achievements in technology must lead to cultural change. Look at how popular vinyl has become, why not theatre again.

I agree, seeing real human actors on stage will always be popular for some consumers. Same for local live musicians. That said, I helped a friend who makes low budget, edgy and cool films last week. I showed him what I knew about driving Pika.art and he picked it up quickly. He is very excited about the possibility of being able to write more stories and turn them into films. I think there is plenty of demand for all…

We will soon find that story generation is easily automated.

Re: Sora: Creating video from text

#809

Obviously incredibly cool, but it seems that people are incredibly overstating the applications of this. Realistically, how do you fit this into a movie, a TV show, or a game? You write a text prompt, get a scene, and then everything is gone—the characters, props, rooms, buildings, environments, etc. won’t carry over to the next prompt.

Nah just fine-tune the model to a specific set of characters or aesthetic. It's not hard, already done with SDXL LoRAs. You can definitely generate a whole movie from just a storyboard.. if not now, then in maybe five yrs.

Re: Sora: Creating video from text

#810
post #141

People here seem mostly impressed by the high resolution of these examples. Based on my experience doing research on Stable Diffusion, scaling up the resolution is the conceptually easy part that only requires larger models and more high-resolution training data. The hard part is semantic alignment with the prompt. Attempts to scale Stable Diffusion, like SDXL, have resulted only in marginally better prompt understan…

[deleted]
Post reply on HN