Live data from Hacker News

Dall-E 2

openai.com

81–90 of 511 posts

Re: Dall-E 2

#81
post #3

Preventing Harmful Generations We’ve limited the ability for DALL·E 2 to generate violent, hate, or adult images. By removing the most explicit content from the training data, we minimized DALL·E 2’s exposure to these concepts. We also used advanced techniques to prevent photorealistic generations of real individuals’ faces, including those of public figures. "And we've also closed off a huge range of potentially int…

This is a horrible idea. So Francis Bacon's art or Toyohara Kunichika's art are out of question.

But at least we can get another billion of meme-d comics with apes wearing sunglasses, so that's good news right?

It's just soul-crushing that all the modern, brilliant engineering is driven by abysmal, not even high-school art-class grade aesthetics and crowd-pleasing ethics that are built around the idea of not disturbing some 1000 very vocal twitter users.

Death of culture really.

Re: Dall-E 2

#82
post #59

I'm only part way through the paper, but what struck me as interesting so far is this: In other text-to-image algorithms I'm familiar with (the ones you'll typically see passed around as colab notebooks that people post outputs from on Twitter), the basic idea is to encode the text, and then try to make an image that maximally matches that text encoding. But this maximization often leads to artifacts - if you ask for…

This isn't something I'm knowledgeable on so forgive my simplification but is this like a sort of micro services for AI. Each AI takes their turn handing some aspect, another sort of mediates among them?

Re: Dall-E 2

#83

Something about this makes me nauseous. Perhaps is the fact that soon the market value for creatives is going to fall to a hair about zero for all but the most famous. We will be all the poorer for it when 95% of images you see are AI generated. There will be niches of course but in a few short years it'll be over for a huge swathe of creative professionals who are already struggling. Some of the images also hit me w…

Not exactly. All the ideas put forth in these demos are really arbitrary, with nothing whatsoever to say. Generating crap art becomes more and more effortless: we've seen this in music as well. Jumping out of the conceptual box to generate novel PURPOSE is not the domain of a Dall-E 2. You've still gotta ask it for things. It's a paintbrush. Without a coherent story, it's an increasingly impressive stunt (or a form o…

This reminds me of an art class in high school in the early 2000s where I handed in a printout of a 3d generated image (painstakingly modeled and rendered in software over the whole weekend by me) and the teacher looked at me and told me that's not art because it's "computer generated" and I didn't "even use my hands" to make it. Even as a teenager, the idea that art is defined by how it's made versus it being a way for the artist to express intention in whatever way they seem fit seemed really reductionist and almost vulgar to me.

Maybe lots artists of the future will actually use AI models to express their inner thoughts and desires in a way that touches something in their audience. It will still be art.

Re: Dall-E 2

#84
post #3

Preventing Harmful Generations We’ve limited the ability for DALL·E 2 to generate violent, hate, or adult images. By removing the most explicit content from the training data, we minimized DALL·E 2’s exposure to these concepts. We also used advanced techniques to prevent photorealistic generations of real individuals’ faces, including those of public figures. "And we've also closed off a huge range of potentially int…

I never considered that our AI overlord could be a prude.

Re: Dall-E 2

#86
post #30

So what does the future of human creativity look like when an AI can generate possibly infinite variations of an idea.

I seem to recall an XKCD that I cannot find, but the premise goes like: When you have a digital display of pixels, if you randomly color pixels at 24 fps then you will eventually display every movie that can be or will ever be made, powerset notwithstanding. This can also be tied to digital audio. In short, while mind-blowingly large, the space of display through digital means is finite.

Sounds a bit like the tower of babel of jorge borges. I imagine most of the videos would be complete random nonsense.

I think an AI infused future is going to become increasingly more absurd and surreal, it will lead to a kind of creative and cultural nihilism, if that's the right term.

Like the value of originality will become meaningless.

Re: Dall-E 2

#87
post #3

Preventing Harmful Generations We’ve limited the ability for DALL·E 2 to generate violent, hate, or adult images. By removing the most explicit content from the training data, we minimized DALL·E 2’s exposure to these concepts. We also used advanced techniques to prevent photorealistic generations of real individuals’ faces, including those of public figures. "And we've also closed off a huge range of potentially int…

Or, it’s a demonstration that AI output can be controlled in meaningful ways, period. Surely this supports openai’s stated goal of making safe AI?

Re: Dall-E 2

#88

Earlier quoted context omitted.

Oh, no, the society! A picture of Joe Biden killing a priest! Society didn't collapse after photoshop. "Responsibility to society" is such a catch-all excuse.

No. Russian society is pretty much collapsing right now under the weight of lies. Currently they are using "it's a fake" to deny their war crimes. Cheap and plentiful is substantivly different from "possible". See for example, oxycontin.

Russia has.. a history of denying the obvious. I come from an ex-communist satellite state so I would know. The majority of the people know what's happening. There's a rather new joke from COVID: the Russians do not take Moderna because Putin says not to trust it, and they do not take Sputnik because Putin says to trust it.

Do not be deluded that our own governments are not manufacturing the narrative too. The US has committed just as many war crimes as Russia. Of course, people feel differently about blowing up hospitals in Afghanistan rather than Ukraine. What the Afghan people think about that is not considered too much.

Re: Dall-E 2

#89
post #10

A friend of mine was studying graphic design, but became disillusioned and decided to switch to frontend programming after he graduated. His thesis advisor said he should be cautious, because automation/AI will soon take the jobs of programmers, implying that graphic design is a safer bet in this regard. Looks like his advisor is a few years from being proven horribly wrong.

Did coachman immediately retire when cars were invented or did they begin personal drivers or taxi drivers?

Re: Dall-E 2

#90
post #73
post #59

I'm only part way through the paper, but what struck me as interesting so far is this: In other text-to-image algorithms I'm familiar with (the ones you'll typically see passed around as colab notebooks that people post outputs from on Twitter), the basic idea is to encode the text, and then try to make an image that maximally matches that text encoding. But this maximization often leads to artifacts - if you ask for…

What always bother me with this stuff is, well, you say one approach is more sensible than the other because the images happen to come out more pleasing. But there's no real rhyme or reason, it is a sort of alchemy. Is text encoding strictly worse or is it an artifact of the implementation? And if it is strictly worse, which is probably the case, why specifically? What is actually going on here? I can't argue that th…

Yeah, I mean you're right that ultimately the proof is in the pudding.

But I do think we could have guessed that this sort of approach would be better (at least at a high level - I'm not claiming I could have predicted all the technical details!). The previous approaches were sort of the best that people could do without access to the training data and resources - you had a pretrained CLIP encoder that could tell you how well a text caption and an image matched, and you had a pretrained image generator (GAN, diffusion model, whatever), and it was just a matter of trying to force the generator to output something that CLIP thought looked like the caption. You'd basically do gradient ascent to make the image look more and more and more like the text prompt (all the while trying to balance the need to still look like a realistic image). Just from an algorithm aesthetics perspective, it was very much a duct tape and chicken wire approach.

The analogy I would give is if you gave a three-year-old some paints, and they made an image and showed it to you, and you had to say, "this looks like a little like a sunset" or "this looks a lot like a sunset". They would keep going back and adjusting their painting, and you'd keep giving feedback, and eventually you'd get something that looks like a sunset. But it'd be better, if you could manage it, to just teach the three-year-old how to paint, rather than have this brute force process.

Obviously the real challenge here is "well how do you teach a three-year-old how to paint?" - and I think you're right that that question still has a lot of alchemy to it.

Post reply on HN