Live data from Hacker News

Sora: Creating video from text

openai.com

251–260 of 1001 posts

Re: Sora: Creating video from text

#251
post #24

Yeah, you just can't let all media, all the cost and hard work of millions of photographers, animators, filmmakers, etc be completely consumed and devalued by one company just because it's a very cool technical trick. The more powerful these services become the more obvious that will be. What OpenAI does is amazing, but they obviously cannot be allowed to capture the value of every piece of media ever created — it'll…

> just because it's a very cool technical trick

That's one big trick, almost magical.

Re: Sora: Creating video from text

#253
Absolutely insane. It's very odd where the glitches happen. Did anyone else notice in the "stylish woman ... Tokyo" clip how her legs skip-hop and then cross at 0:30 in a physically impossible way. Everything else about the clip seems so realistic, yet this is where it trips up?

Re: Sora: Creating video from text

#254
Holy cow, I've literally only looked at the first two videos so far, and it's clear that this absolutely blows every other generative video model out of the water, barely even worth comparing. We immediately jumped from interesting toy models where it was pretty easy to tell that the output was AI generated to.. this.

Re: Sora: Creating video from text

#255

Obviously incredibly cool, but it seems that people are incredibly overstating the applications of this. Realistically, how do you fit this into a movie, a TV show, or a game? You write a text prompt, get a scene, and then everything is gone—the characters, props, rooms, buildings, environments, etc. won’t carry over to the next prompt.

It also seems hard to control exactly what you get. Like you'd want a specific pan, focus etc. to realize your vision. The examples here look good, but they aren't very specific.

But it was the same with Dall-E and others in the beginning, and there's now lots of ways to control image generators. Same will probably happen here. This was a huge leap just in how coherent the frames are.

Re: Sora: Creating video from text

#256

This is insane. But I'm impressed most of all by the quality of motion . I've quite simply never seen convincing computer-generated motion before . Just look at the way the wooly mammoths connect with the ground, and their lumbering mass feels real. Motion-capture works fine because that's real motion, but every time people try to animate humans and animals, even in big-budget CGI movies, it's always ultimately obvio…

> Motion-capture works fine because that's real motion

Except in games where they mo-cap at a frame rate less than what it will be rendered at and just interpolate between mo-cap samples, which makes snappy movements turn into smooth movements and motions end up in the uncanny valley.

It's especially noticeable when a character is talking and makes a "P" sound. In a "P", your lips basically "pop" open. But if the motion is smoothed out, it gives the lips the look of making an "mm" sound. The lips of someone saying "post" looks like "most".

At 30 fps, it's unnoticeable. At 144 fps, it's jarring once you see it and can't unsee it.

Re: Sora: Creating video from text

#257
Many might miss the key paragraph at the end:

   "Sora serves as a foundation for models that can understand and simulate the real world, a capability we believe will be an important milestone for achieving AGI."
This also helps explain why the model is so good since it is trained to simulate the real world, as opposed to imitate the pixels.

More importantly, its capabilities suggest AGI and general robotics could be closer than many think (even though some key weaknesses remain and further improvements are necessary before the goal is reached.)

EDIT: I just saw this relevant comment by an expert at Nvidia:

   “If you think OpenAI Sora is a creative toy like DALLE, ... think again. Sora is a data-driven physics engine. It is a simulation of many worlds, real or fantastical. The simulator learns intricate rendering, "intuitive" physics, long-horizon reasoning, and semantic grounding, all by some denoising and gradient maths.

   I won't be surprised if Sora is trained on lots of synthetic data using Unreal Engine 5. It has to be!

   Let's breakdown the following video. Prompt: "Photorealistic closeup video of two pirate ships battling each other as they sail inside a cup of coffee." ….”
https://twitter.com/DrJimFan/status/1758210245799920123

Re: Sora: Creating video from text

#258

Just in time for the election season. Also "A cat waking up its sleeping owner demanding breakfast" has too many paws - yes I do feel petty saying this.

And the sleeper's shoulder gets converted to the duvet? And a strange extra hand somewhere. It was also the one that to me stood out as the worst. The quality was good, but it had the same artifacts as previous generations of ai videoes where thing morphs.

Re: Sora: Creating video from text

#260

I am a CG artist and Director and this made me so sad. I am watching in horror and amazement. I am not anti AI at all, but being on the wrong side of efficiency, for the individual this is heartbreaking. its so much fun to make CG and create shots and the reason its hard (just like anything) makes it rewarding.

I'm conflicted though because on the flip side it could open up filmmaking to way more people who don't have the skills/money/time

Like what if any artist could make a whole movie by themself without needing millions of dollars or hundreds of people

Similar to how you used to need a huge studio full of equipment to record music and now someone in their bedroom with a DAW can do it

Post reply on HN