Live data from Hacker News

Sora is here

openai.com

171–180 of 1001 posts

Re: Sora is here

#171
post #124
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…

For those not in this space, Sora is essentially dead on arrival.

Sora performs worse than closed source Kling and Hailuo, but more importantly, it's already trumped by open source too.

Tencent is releasing a fully open source Hunyuan model [1] that is better than all of the SOTA closed source models. Lightricks has their open source LTX model and Genmo is pushing Mochi as open source. Black Forest Labs is working on video too.

Sora will fall into the same pit that Dall-E did. SaaS doesn't work for artists, and open source always trumps closed source models.

Artists want to fine tune their models, add them to ComfyUI workflows, and use ControlNets to precision control the outputs.

Images are now almost 100% Flux and Stable Diffusion, and video will soon be 100% Hunyuan and LTX.

Sora doesn't have much market apart from name recognition at this point. It's just another inflexible closed source model like Runway or Pika. Open source has caught up with state of the art and is pushing past it.

[1] https://github.com/Tencent/HunyuanVideo

Re: Sora is here

#172
post #158
post #114

Earlier quoted context omitted.

That limits its value for industries like Hollywood, though, doesn't it? And without that, who exactly is going to pay for this?

Advertisers, I guess. Same folks who paid for everything else around here

Yeah, I just question if there are enough customers to make this work.

Re: Sora is here

#173
post #124

Earlier quoted context omitted.

It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…

something like a white paper with a mood board, color scheme, and concept art as the input might work. This could be sent into an LLM "expander" that increases the words and speficity. Then multiple reviews to tap things in the right direction.

And I think this realistically is going to be the shape of the tools to come in the foreseeable future.

Re: Sora is here

#174

Hollywood's days are numbered. If you are a creative in this industry, start preparing to transition to another industry or adapt. Your boss is highly likely to be toying around with this. The first entirely AI generated film (with Sora or other AI video tools) to win an Oscar will be less than 5 years away.

Nothing I'm seeing here looks like it's going to destroy Hollywood. I could see this tool maybe being used for generating establishing shots (generate a sweeping drone shot of a lighthouse looking out over a stormy sea), but then the actual talent work in a scene will be way more sensitive. The little details matter so much, and this feels so far from getting all of that right. Sure, this is the worst it will ever be…

I'm not sure the little details are enough of a moat. Consider TikTok - people use cheap "special effects" to get the message across, e.g. if a man is playing a woman he might drape a towel over his head - it's silly and low quality but it gets the idea across to the viewer. Think too about programs like Archer or South Park that have (stylistically) low quality animation but still huge fan bases.

What I think this will unlock, maybe with a bit of improvement, is low quality video generation for a vast number of people. Do you have a short film idea? Know people with some? Likely millions of people will be able to use this to put together good enough short films - that yes, have terrible details, but are still good enough to watch. Some of those millions of newly enabled videos will have such strong ideas or writing behind them that it will make up for, or capitalize on, the weak video generation.

As the tools become easier, cheaper, faster, better etc more and more hobbyists will pick them up and try to use them. The user base will encourage the product to grow, and it will gradually consume film (assuming it can reach the point of being as or nearly as good as modern special effects).

I think of it like - when Steven Spielberg was young he used an 8mm camera, not as good as professional film equipment in the day, but good enough to create with. If I were a high school student interested in film I would absolutely be using stuff like this to create.

Re: Sora is here

#175
post #124
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…

> Under the hood, with the way the text is turned into vector embeddings, it's fairly questionable whether you'd agree that it can even represent such a thing.

The text encoder may not be able to know complex relationships, but the generative image/video models that are conditioned on said text embeddings absolutely can.

Flux, for example, uses the very old T5 model for text encoding, but image generations from it can (loosely) adhere to all rules and nuances in a multi-paragraph prompt: https://x.com/minimaxir/status/1820512770351411268

Re: Sora is here

#176
“Right before the TikTok ban goes into effect” is incredible market timing for the release of a tool that is useless for anything other than terrible TikTok spam videos

Re: Sora is here

#177
post #162

Anyone else find this stuff extremely distasteful? "Disrupting" creativity and art feels like it goes against our humanity.

It is like an attempt to do psychic battle over the meaning of "disruption".

Re: Sora is here

#178
post #17

Not available in the EU: https://help.openai.com/en/articles/10250692-sora-supported-...

Does VPN solves the problem? I'm living in an EU country and I don't like that the EU decides for me (and companies like OpenAI or Meta don't give out their models to me)! I'm an old enough adult to decide for myself what I want...

I used a Japanese protonVPN , I got past the "Not in the EU" thing but it said "no new signups are allowed atm".

Perhaps just best to wait

Re: Sora is here

#179

Wow this is bad. And by bad i mean worse than leading open source and existing alternatives. Is it me or does it seem like OpenAI revolutionized with both chatGPT and Sora, but they've completely hit the ceiling? Honestly a bit surprised it happened so fast!

What are the leading alternatives? (Open source or otherwise)

FLUX

Re: Sora is here

#180
post #9

I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…

> A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch.

Sure, if you then do the same in reverse.

Post reply on HN