I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…
It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…
Sora performs worse than closed source Kling and Hailuo, but more importantly, it's already trumped by open source too.
Tencent is releasing a fully open source Hunyuan model [1] that is better than all of the SOTA closed source models. Lightricks has their open source LTX model and Genmo is pushing Mochi as open source. Black Forest Labs is working on video too.
Sora will fall into the same pit that Dall-E did. SaaS doesn't work for artists, and open source always trumps closed source models.
Artists want to fine tune their models, add them to ComfyUI workflows, and use ControlNets to precision control the outputs.
Images are now almost 100% Flux and Stable Diffusion, and video will soon be 100% Hunyuan and LTX.
Sora doesn't have much market apart from name recognition at this point. It's just another inflexible closed source model like Runway or Pika. Open source has caught up with state of the art and is pushing past it.