Earlier quoted context omitted.
> If you haven't been consuming everything on the internet with a high alert bs sensor, then that's an issue of its own "just be privileged as I was to get all the necessary education to be able to not be fooled by this tech". Yeah, very realistic and compassionate.
What education do you specifically think is necessary for people with average IQs all over the world to not be fooled by this, given that they are aware that videos can easily be faked in 2024? A high school degree? A bachelors?
Sora is here
761–770 of 1001 posts
Re: Sora is here
#762https://app.checkbin.dev/snapshots/1f0f3ce3-6a30-4c1a-870e-2...
Pros:
- Some of the Sora results are absolutely stunning. Check out the detail on the lion, for example! - The landscapes and aerial shots are absolutely incredible. - Quality is much better than Mochi & LTX out of the box. Mochi/LTX seem to require specifically optimized workflows (I've seen great img2vid LTX results on Reddit that start with Flux image generations, for example). Hunyuan seems comparable to Sora!
Cons:
- Still nearly impossible to access Sora despite the “launch”. My generations today were in the 2000s, implying that it’s only open to a very small number of people. There’s no api yet, so it’s not an option for developers. - Sora struggles with physical interactions. Watch the dancers moonwalk, or the ball goes through the dog. HunyuanVideo seems to be a bit better in this regard. - Can't run it locally mode (obviously) - I haven't tested this, but I think it's safe to assume Sora will be censored extensively. HunyuanVideo is surprisingly open (I've seen NSFW generations!) - I’m getting weird camera angles from Sora, but that could likely be solved with better prompting.
Overall, I’d say it’s the best model I've played with, though I haven’t spent much time on other non-open-source ones. Hunyuan gives it a run for its money, though!
Re: Sora is here
#763So in the end, was it all classic hype pumping and heavily edited marketing material? Can't create account but from what I see on the front-page, every video looks kinda strange, almost in everyone of them something is instantly off to my eye. Anyone got decent results yet?
As far as I can tell, this is not Sora, but a distilled model that runs in a reasonable amount of time, with a reasonable amount of compute. It's pretty likely that's resulted in degradation of quality. Further, the marketing/demo-ing for Sora for the past year has been heavily curated videos from OpenAI and it's not clear what was generate using Sora and what was being generated using Sora "Turbo" (this distilled mo…
Re: Sora is here
#764Earlier quoted context omitted.
As far as I can tell, this is not Sora, but a distilled model that runs in a reasonable amount of time, with a reasonable amount of compute. It's pretty likely that's resulted in degradation of quality. Further, the marketing/demo-ing for Sora for the past year has been heavily curated videos from OpenAI and it's not clear what was generate using Sora and what was being generated using Sora "Turbo" (this distilled mo…
Read few other reviews as well, the general feeling seems to be more or less the same. People also complain that it often imagines faces on reference pictures and after some substantial delay denies generating, which is a big game-ender.
Hopefully these types of issues blow over as they increase capacity or load decreases.
The lengthy generation times aren't fun to deal with though in any case. As good as the UX for the app itself is, there's little they can do about how long it takes for a video to generate compared to images. The near instant feedback is gone (just like old times)
Re: Sora is here
#765For those curious (and still locked out) here’s direct a comparison of Sora vs. the open-source leaders (HunyuanVideo, Mochi and LTX): https://app.checkbin.dev/snapshots/1f0f3ce3-6a30-4c1a-870e-2... Pros: - Some of the Sora results are absolutely stunning. Check out the detail on the lion, for example! - The landscapes and aerial shots are absolutely incredible. - Quality is much better than Mochi & LTX out of the bo…
Re: Sora is here
#766I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…
It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…
Re: Sora is here
#767I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…
It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…
Re: Sora is here
#768Earlier quoted context omitted.
a somewhat counterintuitive argument is this: AI models will make the overall creative landscape more diverse and interesting, ie, less "average"! Imagine the space of ideas as a circle, with stuff in the middle being more easy to reach (the "cliched average"). Previously, traversing the circle was incredibly hard - we had to use tools like DeviantArt, Instragram, etc to agglomerate the diverse tastes of artists, hop…
To extend the analogy, imagine the circle as a probability distribution; for simplicity, imagine it's a bivariate normal joint distribution (aka. Gaussian in 3D) + some noise, and you're above it and looking down. When you're commissioning an artist to make you some art, you're basically sampling from the entire distribution. Stuff in the middle is, as you say, easiest to reach, so that's what you'll most likely get.…
Tastes are almost never normally distributed along a spectrum, but multi-modal. So the more dimensions you explore in, the more you end up with “islands of taste” on the surface of a hyper sphere and nothing like the normal distribution at all. This phenomenon is deeply tied to why “design by committee” (eg, in movies) always makes financial estimates happy but flops with audiences — there is almost no customer for average anything.
I agree with your conclusion.
Re: Sora is here
#769I've found using these and similar tools that the amount of prompts and iteration required to create my vision (image or video in my mind) is very large and often is not able to create what I had originally wanted. A way to test this is to take a piece of footage or an image which is the ground truth, and test how much prompting and editing it takes to get the same or similar ground truth starting from scratch. It is…
It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…
Re: Sora is here
#770For those curious (and still locked out) here’s direct a comparison of Sora vs. the open-source leaders (HunyuanVideo, Mochi and LTX): https://app.checkbin.dev/snapshots/1f0f3ce3-6a30-4c1a-870e-2... Pros: - Some of the Sora results are absolutely stunning. Check out the detail on the lion, for example! - The landscapes and aerial shots are absolutely incredible. - Quality is much better than Mochi & LTX out of the bo…
The vibe they give me is similar to the iPhone photography commercials where yes, in theory, a picnic in the park could look exactly like this except for all the parts that seem movie perfect.
I guess it's really more of a colour grading question where most of the Sora colour grading triggers that part of my brain that says "I'm watching a movie and this isn't real" without quite realising why.
A few of the Hunyuan videos in contrast seem a bit more believable even though they have some obvious glitches at times.
The other thing I think Sora has is that thing in commercials where no one else except the protagonist exists and nothing is ever inconvenient. The video of the teacher in a classroom with no students reminds me of that as well as the picnic in the park where there's wide open space with no one around.
I suppose it depends if the goal is to generate believable video and how you define believable.