The concern trolling and gatekeeping about social justice issues coming from the so-called "ethicists" in the AI peanut gallery has been utterly ridiculous. Google claims they don't want to release Imagen because it lacks what can only be called "latent space affirmative action".
Stability or someone like it will valiantly release this technology, again and there will be absolutely no harm to anyone.
Stop being so totally silly Google, OpenAI, et. al. - it's especially disingenuous because the real reason you don't want to release these things is that you can't be bothered to share and would rather keep/monetize the IP. Which is ok -- but at least be honest.
Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…
Why did you create a throwaway to post this? I've seen a lot of Stable Diffusion promoters on various platforms recently, with similarly new accounts. What is up with that?
It's quite simply because I'm on my work computer, and I wanted to fire off a comment here. No nefarious purposes. My regular account is uejfiweun.
The concern trolling and gatekeeping about social justice issues coming from the so-called "ethicists" in the AI peanut gallery has been utterly ridiculous. Google claims they don't want to release Imagen because it lacks what can only be called "latent space affirmative action". Stability or someone like it will valiantly release this technology, again and there will be absolutely no harm to anyone. Stop being so to…
I agree basically completely, but there’s now a cottage industry of AI Ethics professionals whose real job is to provide a smoke screen for the “cake and eat it too” that the big shops want on this kit: peer review and open source contributions and an academic atmosphere when it suits them, proprietary when it doesn’t. Those folks are a lobby now.
The thing about owning the data sets and the huge TPU/A100 clusters is that the “publish the papers” model strictly serves them: no one can implement their models, they can implement everyone else’s.
It's interesting that these models can generate seemingly anything, but the prompt is taken only as a vague suggestion. From the first 15 examples shown to me, only one contained all elements of the prompt, and it was one of the simplest ("an astronaut riding a horse", versus e.g. "a glass ball falling in water" where it's clear it was a water droplet falling and not a glass ball). We're seeing leaps in random capabi…
The way I see it, input being confined to a "text description" is the next immediate problem that needs to be solved. I don't think we can rely on textual inputs for much longer as the human language is too imprecise and/or verbose. It's hard to imagine what exactly the optimal interface would be, but I'm thinking we'll need ways to dictate attributes for each entity being represented, the backdrop, and the view composition all as separate individual components. Ideally all these components can also be reusable and provide reproducibility guarantees without having to share a global "seed" as well.
How long until the AI just generates the entire frame buffer on a device? Then you don’t need to design or program anything; the AI just handles all input and output dynamically.
The concern trolling and gatekeeping about social justice issues coming from the so-called "ethicists" in the AI peanut gallery has been utterly ridiculous. Google claims they don't want to release Imagen because it lacks what can only be called "latent space affirmative action". Stability or someone like it will valiantly release this technology, again and there will be absolutely no harm to anyone. Stop being so to…
There is a clear risk from these sorts of models as they get better - I mean recreating specific individuals’ likenesses in compromising images (or even worse, video). We’re not at that point yet, but these things are getting better fast, so it’s only a matter of time. The problem is that there’s no way to mitigate those risks except to keep the model behind an inference-only API, or not release it at all - as soon as the model is open sourced then fine-tuning the model can introduce whatever behavior you want. Holding the models back is virtue signaling at best and actively harmful at worst, because it draws attention from individuals who will take it as a challenge to find ways to misuse them, and cuts businesses and the open source community off from models that would otherwise be very useful to them, that they could help to improve. I’m concerned this is becoming a self fulfilling prophecy, where more companies will start to do the same thing, primarily because it’s what everyone else is doing.
How long until the AI just generates the entire frame buffer on a device? Then you don’t need to design or program anything; the AI just handles all input and output dynamically.
Imagine you click a youtube video in a bad network envoirment, then the server sends like an alt tag equivalent for the video as a promnt, and the Neural Engine chip inside your phone create the first seconds of the video while it loads.
We're fay away from it now, but I've seen less sketchy solutions being implemented.
Emad (founder of Stability AI) has said they already have video model training underway, as well as text and audio. Exciting times.
Is this going to end up into a single model, where its trained on text and images and audio and videos and 3d models, and it can do anything to anything depending on what you ask of it? Feels like the cross-training would help yield stronger results.
It's interesting that these models can generate seemingly anything, but the prompt is taken only as a vague suggestion. From the first 15 examples shown to me, only one contained all elements of the prompt, and it was one of the simplest ("an astronaut riding a horse", versus e.g. "a glass ball falling in water" where it's clear it was a water droplet falling and not a glass ball). We're seeing leaps in random capabi…
The way I see it, input being confined to a "text description" is the next immediate problem that needs to be solved. I don't think we can rely on textual inputs for much longer as the human language is too imprecise and/or verbose. It's hard to imagine what exactly the optimal interface would be, but I'm thinking we'll need ways to dictate attributes for each entity being represented, the backdrop, and the view comp…
Like an AI-assisted photoshop, but that's not restricted by language, only interactivity. Down the line you'll need a direct mind meld because some ideas don't have words to describe them, but that's not the "next immediate" problem.
I’m going to post an Ask HN about what am I supposed to do when I’m “disrupted”. I work in film / video / CG where the bread and butter is short form advertising for Youtube, Instagram and TV. It’s painfully obvious that in 1 year the job might be exceedingly more difficult than it is now.
the principle of least action says you will move to adjacent territory. either you become and advertiser, or you learn to make these models