Live data from Hacker News

DALL-E 2 has a secret language

twitter.com

11–20 of 118 posts

Re: DALL-E 2 has a secret language

#11
post #7

Possibly related: In 2017 AI bots formed a derived shorthand that allowed them to communicate faster: https://www.facebook.com/dhruv.batra.dbatra/posts/1943791229... > While the idea of AI agents inventing their own language may sound alarming/unexpected to people outside the field, it is a well-established sub-field of AI, with publications dating back decades. > Simply put, agents in environments attempting to solv…

Unintuitive to biased humans. The solutions may actually be super intuitive/efficient, and we just can't wrap our heads around it yet

Re: DALL-E 2 has a secret language

#12
It seems obvious this would happen (it's just adversarial inputs again) - they didn't make DALL-E reject "nonsense" prompts, so it doesn't try to, and indeed there's no reason you'd want to make it do that.

Seems like a useful enhancement would be to invert the text and image prior stages, so it'd be able to explain what it thinks your prompt meant along with making images of it.

Re: DALL-E 2 has a secret language

#14

Interestingly Google detects these words as Greek. I know they are nonsensical and not actually Greek but I'm wondering if any Greek speakers might be able to provide some insights. Are these gibberish words close to meaningful words? (clear shot in the dark here) Maybe a linguist could find more meaning?

One could conjecture that "Apoploe" is similar to από πουλί, "from bird". But I don't have much support for that conjecture.

Re: DALL-E 2 has a secret language

#15

It’s wild to see the discoveries being made in ML research. Like most of these ‘discoveries,’ it makes a fair amount of sense after thinking about it. Of course it’s not just going to spit out random noise for random input, it’s been trained to generate realistic looking images. But I think it is an interesting discovery because I don’t think anyone could have predicted this. One of my favorite examples is the classi…

> One of my favorite examples is the classification model that will identify an apple with a sticker on it that says “pear” as a pear—it makes sense, but is still surprising when you first see it.

That classification model (CLIP) is the first stage of this image generator (DALLE) - and actually this shows that it doesn't think they're exactly the same thing, or at least that's not the full story, because DALL-E doesn't confuse the two.

However, other CLIP guided image generation models do like to start writing the prompt as text into the image if you push them too hard.

Re: DALL-E 2 has a secret language

#18

It seems obvious this would happen (it's just adversarial inputs again) - they didn't make DALL-E reject "nonsense" prompts, so it doesn't try to, and indeed there's no reason you'd want to make it do that. Seems like a useful enhancement would be to invert the text and image prior stages, so it'd be able to explain what it thinks your prompt meant along with making images of it.

[deleted]
Post reply on HN