Earlier quoted context omitted.
I wish Google would allow me to remove the AI stuff from search results. 99% of the times it's either useless or wrong.
Add a -ai to the end of your Google search query. There are also browser extensions that stop the AI content from displaying. I use the one for Chrome called "Remove Google Search Generative AI".
Sora is here
661–670 of 1001 posts
Re: Sora is here
#662Re: Sora is here
#663Earlier quoted context omitted.
> I don't think it'll be possible for a closed source tool to compete with the open image/video ecosystem. And I don't think the current status quo of open source models being entirely subsidised by startups and corporations is sustainable, they're all hemorrhaging money and their investors will only have so much patience before they expect returns. Enjoy it while it lasts.
It's game theory. If you don't have market share for your closed model, you release it as open source and let a community build upon it. Mochi is better positioned to build tools on top of their community model. They're already thinking about control. Weights are commodity. Products have value.
Stability was supposed to be doing a similar "give away the models but sell products built on them" strategy and it doesn't seem to be working for them, by all accounts they're barely able to keep the lights on.
Re: Sora is here
#664Earlier quoted context omitted.
There's big difference between cartoonishly incorrect and uncanny valley plausibly correct.
There's a huge amount of such stuff in movies. Special effects, weapons physics, unrealistic vehicles and planes, or the classic 'hacking'.
Re: Sora is here
#665Earlier quoted context omitted.
> Long term, you'll never have a coherent movie produced by stringing together a series of textual snippets because, again, that's just impossible. Why snippets? Submit a whole script the way a writer delivers a movie to a director. The (automated) director/DP/editor could maintain internal visual coherence, while the script drives the story coherence.
You should watch how movies are made sometime. How a script is developed. How changes to it are made. How storyboards are created. How actors are screened for roles. How locations are scouted, booked, and changed. How the gazillion of different departments end up affecting how a movie looks, is produced, made, and in which direction it goes (the wardrobe alone, and its availability and deadlines will have a huge impa…
At the same time I am curious in the "that person has too many fingers" sense at what a system trained on tens of thousands of movies plus scripts plus subtitles plus metadata etc. would generate.
I thought about it for a bit and I would want to watch a computer generated Sharknado 7 or Hallmark Christmas movie.
Re: Sora is here
#666Re: Sora is here
#667Earlier quoted context omitted.
It just plain isn't possible if you mean a prompt the size of what most people have been using lately, in the couple hundred character range. By sheer information theory, the number of possible interpretations of "a zoom in on a happy dog catching a frisbee" means that you can not match a particular clip out of the set with just that much text. You will need vastly more content; information about the breed, informati…
> Under the hood, with the way the text is turned into vector embeddings, it's fairly questionable whether you'd agree that it can even represent such a thing. The text encoder may not be able to know complex relationships, but the generative image/video models that are conditioned on said text embeddings absolutely can. Flux, for example, uses the very old T5 model for text encoding, but image generations from it ca…
Flux certainly does not consistently do so across an arbitrary collection of multi-paragraph prompts, as anyone whose run more than a few long prompts past it would recongize; also, the tweet is wrong in the other direction, as well, longer language-model-preprocessed prompts for models that use CLIP (like various SD1.5 and SDXL derivatives) are, in fact, a common and useful technique. (You’d kind of think that the fact that generated prompt here is significantly longer than the 256 token window of T5 would be a clue that the 77 token limit of CLIP might not be as big of a constraint as the tweet was selling it as, too.)
Re: Sora is here
#668Earlier quoted context omitted.
With a heavy dose of "if masses of people are fooled by this, it can't affect me as long as I can see through it. No possible repercussions of mass people believing completely made up stuff that could affect laws, etc."
This entire thread reeks of "I'm smart enough to know that videos can be faked, but Jethro in the trailer park isn't because he's just a plumber, and therefore this tech needs to be censored or else Jethro might believe stuff that makes him vote in a way I don't like" going on here. While the average person overestimates their own intelligence, the average techy dramatically underestimates the intelligence of the ave…
Example: if a gen ai vid of a politician doing some crazy crime came out. Even if it were proven fake, people would start questioning everything and still act as if the politician were guilty
Re: Sora is here
#669Wow this is bad. And by bad i mean worse than leading open source and existing alternatives. Is it me or does it seem like OpenAI revolutionized with both chatGPT and Sora, but they've completely hit the ceiling? Honestly a bit surprised it happened so fast!
Each company would either rush to get a phone out with the new snapdragon chip, or take their time to polish a release and have a better phone late cycle. But the real improvements we're just the chip.
Nvidia chips/larger data centers are the chips. the models are the plethora of android phones each generation.
That kept going until progress stabilized. Then the best user experience & vertical integration won over chasing chip performance (apple).
Re: Sora is here
#670Earlier quoted context omitted.
This almost certainly won’t work. Feel free to feed any of the hundreds of existing film scripts and test how coherent the models can be. My guess is not at all
The clips on the Sora site today would have been utterly astonishing ten years ago. Long term progress can be surprising.
Yeah, and Apollo 11 would have been utterly astonishing a decade before it occurred. And, yet, if you tried to project out from it to what further frontiers manned spaceflight would reach in the following decades, you’d…probably grossly overestimate what actually occurred.
> Long term progress can be surprising.
Sure, it can be surprising for optimists as well as naysayers; as a good rule of thumb, every curve that looks exponential in an early phase ends up being at best logistic.