Live data from Hacker News

Stable Diffusion Public Release

stability.ai

391–400 of 437 posts

Re: Stable Diffusion Public Release

#391

Earlier quoted context omitted.

I think right now we could setup the AI's to do the prompts. You type in a vague description - "gorilla in a suit" and that is passed to GPT-3's API with instructions to provide a detailed and vivid description of the input in X style where X is one of several different styles. GPT-3 generates multiple prompts, the prompts passed to Stable Diffusion, the user gets back a grid of different images and styles. Selecting…

Stable Diffusion already includes a GPT-type language model - that's what's understanding your text prompt.

Sure it's a similar architecture but c'mon, 307 million learnable params for ViT-L vs 175 billion for GPT-3...

Re: Stable Diffusion Public Release

#392

Played with it for a bit in DreamStudio so I could control more of the settings. So far everything it generates is "high quality", but the AI seems to lack the creativity and breadth of understanding that DALL-E 2 has. OpenAI's model is better at taking wildly differing concepts and figuring out creative ways to glue them together, even if the end result isn't perfect. Stable Diffusion is very resistant to that, and…

And then you have Midjourney which would most likely be the best one at doing that specific request. The most crazily creative artistic of the 3.

MJ has announced a new beta mode, that uses SD under the hood.

Re: Stable Diffusion Public Release

#393
post #208

Earlier quoted context omitted.

Not really. I can't think of a recent "leaked texts" that the participants cannot not easily and plausibly deny (e.g. elon's supposed messages to gates), or even voice messages. Even most images can already be denied as photoshop if all the witnesses agree. The only medium that is somewhat hard to deny is videos, like sex tapes, but that's also not too hard. I think there will soon be a race to make deep learning pic…

It'll be hard to deny crypto-signed photos https://petapixel.com/2022/08/08/sonys-forgery-proof-tech-ad... , especially if they include metadata that distinguish photos of AI generated images from normal photos.

No post body was provided.

Re: Stable Diffusion Public Release

#394
post #118

Earlier quoted context omitted.

I much prefer a company that asks me to not abuse it but lets me, then a company that treats me like child and force filters it out for me.

That's not the only issue with NSFW; the larger problem is when you don't ask for it and get it anyway. Especially because this model is not particularly good at it and you'll get body horror.

Well it comes with an NSFW filter that's activated by default so that's a non-issue (but importantly you can turn it off if you're ok with accidentally seeing stuff like that without explicitly asking for it).

Re: Stable Diffusion Public Release

#395
post #115

Earlier quoted context omitted.

Why would people want to consume art that says nothing and means nothing? While this technology is fascinating, it produces the visual equivalent of muzak, and will continue to do so in perpetuity without the ability to reason.

That's the problem for me too. This tech is cool for games, stock images etc but for actual art it's pretty meaningless. The artist's experience, biography and relationship with the world and how that feeds into their work is the WHOLE point for me. I want to engage, via any real artistic product, with a lived human experience. Human consciousness in other words. To me this technology is very clever but it's meaningl…

Oh, nice take! Reminds me of the chess world, long time ago chess entities that could beat any human in the world have existed, people wouldn't play them because losing all the time was boring.

But then came the neural network chess engines (powered by the same training technology of text to images generators), which even people enjoyed playing for some time due to novelty and how they learned how to play from scratch (Alpha Zero, Lc0, et al.)

In the process to get there, we got networks of all kinds of strengths, you could find one exactly as strong as you, and its mistakes were like the mistakes you would make.

And yet, they were missing the "human factor", people would rather play against other humans online, some even willing to pay for accounts in places like playchess and chess.com to play other humans.

As the networks became stronger and then stockfish assimilated them to get the best of both worlds with NNUE, nobody cared and there was even an explosion of human vs. human chess (when twitch.tv and youtube stars were playing each other and the audience didn't even care how bad the chess was, it turned out that it held up as a great spectacle despite, or thanks to the stars being novices - and the race to get better at chess.)

Now chess bots are just a curiosity and only used by people without online access to other humans, I wonder if it'll be the same for art, and if "show your work" becomes a thing.

You can ask an AI to produce a great picture... now try asking it to make a video of you making that art from scratch and the creative process - the whole art section of twitch tv is about the artist's process, and yes, there'S some people that get enough from that to dedicate all their time to their art.

Re: Stable Diffusion Public Release

#396

Earlier quoted context omitted.

Considering the phenomenal progression of Dall-E-1 to Dall-E-2 in just over a year, I'm not really understanding your confidence on the limits of AI content generation.

An AI would need to generate: * Coherent video * Characters with backstory * Dialogue (including jokes and witty banter) * Music ...among many other things. Plus, the training set for video is orders of magnitude smaller than for digital art. (And is additionally burdened with copyright issues.) As I see it, there's simply no path from the DALL-E of today to something like that. And all for art that, essentially, "sa…

Check out https://github.com/THUDM/CogVideo - progress is being made on coherent video generation.

Characters and dialogue are effectively solved, just look at GPT-3.

The entity behind StableDiffusion is also supporting generative music art, so let's see what is coming out of that: https://www.harmonai.org/

We are currently far away from generating a production quality movie with AI, but I don't think it's going to be nearly as long as a lifetime. In my opinion, we'll have high quality AI shorts within the decade.

Re: Stable Diffusion Public Release

#397

Earlier quoted context omitted.

An AI would need to generate: * Coherent video * Characters with backstory * Dialogue (including jokes and witty banter) * Music ...among many other things. Plus, the training set for video is orders of magnitude smaller than for digital art. (And is additionally burdened with copyright issues.) As I see it, there's simply no path from the DALL-E of today to something like that. And all for art that, essentially, "sa…

Check out https://github.com/THUDM/CogVideo - progress is being made on coherent video generation. Characters and dialogue are effectively solved, just look at GPT-3. The entity behind StableDiffusion is also supporting generative music art, so let's see what is coming out of that: https://www.harmonai.org/ We are currently far away from generating a production quality movie with AI, but I don't think it's going to b…

>Characters and dialogue are effectively solved, just look at GPT-3.

Is this the motherload of exaggeration?

Current language models cannot generate coherent dialog (and even then it's mostly bad dialog) spanning more than a minute or two. And their current capabilities in that area are definitely significantly below those of the average human writer.

Re: Stable Diffusion Public Release

#398

Earlier quoted context omitted.

Considering the phenomenal progression of Dall-E-1 to Dall-E-2 in just over a year, I'm not really understanding your confidence on the limits of AI content generation.

An AI would need to generate: * Coherent video * Characters with backstory * Dialogue (including jokes and witty banter) * Music ...among many other things. Plus, the training set for video is orders of magnitude smaller than for digital art. (And is additionally burdened with copyright issues.) As I see it, there's simply no path from the DALL-E of today to something like that. And all for art that, essentially, "sa…

That's true. But the thing with technology, and the reason we've kept up with Moore's law is that someone eventually has a bright idea that leaves current methods and improvement extrapolation in the dust, and then the real thing happens earlier than the most optimistic dates, and performs better than what people expected. The question is not if one day an AI can generate a movie that you can't differentiate from a human-made movie. The question is how long will it take for an AI generated movie to be better than all the human-made movies in history, if it's possible at all it'll happen much sooner than people think possible.

Re: Stable Diffusion Public Release

#399

Does anyone know why I get a SIGSEGV on Ubuntu on a 6 GB VRAM 1060? I even used the optimized script, which needs less VRAM.

Did you make sure to use the 16 bit float weights?

I used the optimized script from https://github.com/basujindal/stable-diffusion, I think it does that automatically?

Re: Stable Diffusion Public Release

#400

Earlier quoted context omitted.

I have a friend who works as an artist and he's excited and nervous about this. But also he's trying to learn how to use these well. If you try these AI out, there's definitely an art to writing good prompts that gets what you actually want. Hopefully these AI will become just another brush in the artist's palette rather than a replacement. I hope these end up similar to the relationship between Google and programmin…

I have a friend who paints, google/ athanasart. His paintings are usually purchased before they are even finished, and sometimes even stolen. All of that technology, Dalle2, Disco diffusion, Stable Diffusion, is totally boring to him. He doesn't even care. He will not install this technology to his machine, or use it online, even if he is paid to do it. I love Craiyon, Dalle2, Stable Diffusion, but just beautiful ima…

> but just beautiful images, are not art.

But the definition of art can't depend on your knowledge of how was it made. I can show you two beautiful pictures and ask you if any of them is art, and you'd know I was trying to trick you. But you could pass a piece made by an AI as art if you didn't know how it was produced.

Post reply on HN