Live data from Hacker News

DALL-E 2 has a secret language

twitter.com

61–70 of 118 posts

Re: DALL-E 2 has a secret language

#61
post #36
post #16

One of the replies is a thread with a fairly convincing rebuttal, with examples: https://twitter.com/Thomas_Woodside/status/15317102510150819...

I'm not sure it's a convincing rebuttal, the examples shown all seem to have some visible commonality. Eg. "Apoploe vesrreaitais" Could refer to something along the lines of a "fan / wedge" or "wing-like" If you look at the examples of cheese, when compared to the "birds and cheese" the cheese tends to be laid out in a fan like pattern and shaped in sharp angled wedges.

> Apoploe vesrreaitais" Could refer to something along the lines of a "fan / wedge"

"feathered" maybe?

Re: DALL-E 2 has a secret language

#62

Interestingly Google detects these words as Greek. I know they are nonsensical and not actually Greek but I'm wondering if any Greek speakers might be able to provide some insights. Are these gibberish words close to meaningful words? (clear shot in the dark here) Maybe a linguist could find more meaning?

As a native Greek, no, they don't make any sense.. sort of. My hunch is that they read significantly more like Latin than they do Greek. However it tells us something about google translate. The reason "Apoploe vesrreaitais" is detected as Greek is because the first "word" is "phonetically" similar to the word απόπλους, which means sailing/shipping and it is rooted in ancient Greek. If we were to write Αποπλοuς using…

I don't think that's how language detection works, they most likely use the frequencies of n-grams to detect language probability. It's still detected as Greek if you change to "Apoulon vesrreaitais", just because it kind of looks the way Greek words look, not because it resembles any specific word.

Re: DALL-E 2 has a secret language

#65
post #34

Shouldn't this be expected to a certain extent? Gibberish has to map _somewhere_ in the models concept space. Whether is maps onto anything we'd recognise as consistent doesn't mean that the AI wouldn't have some concept of where it relates, as other people have noted, the gibberish breaks down when you move it into another context, but who's to say that Dall-E 2 isn't remaining consistent to some concept it understa…

This is really interesting because I was just looking at gibberish detection using GPT models. Seems like mitigating AI with AI doesn't sound like it's all that secure since you can probably mess with the gibberish detection similarly - Or maybe the 'secret language' as they're calling it here passes GPT gibberish detection? [1]

[1] https://arr.am/2020/07/25/gpt-3-uncertainty-prompts/

Re: DALL-E 2 has a secret language

#66
Wait, how does that make any sense?

I thought DALL-E's language model was tokenized, so it doesn't understand that eg "car" is made up of the letters 'c', 'a' and 'r'.

So how could the generated pictures contain letters that form words that are tokenized into DALL-E's internal "language"? Shouldn't we expect that feeding those words to the model would give the same result as feeding it random invented words?

Actually, now that I think about it, how does DALL-E react when given words made of completely random letters?

Re: DALL-E 2 has a secret language

#67
post #62

Earlier quoted context omitted.

As a native Greek, no, they don't make any sense.. sort of. My hunch is that they read significantly more like Latin than they do Greek. However it tells us something about google translate. The reason "Apoploe vesrreaitais" is detected as Greek is because the first "word" is "phonetically" similar to the word απόπλους, which means sailing/shipping and it is rooted in ancient Greek. If we were to write Αποπλοuς using…

I don't think that's how language detection works, they most likely use the frequencies of n-grams to detect language probability. It's still detected as Greek if you change to "Apoulon vesrreaitais", just because it kind of looks the way Greek words look, not because it resembles any specific word.

You are wrong. Had it been that simple I would __not__ have suggested that and for whatever reason I find your reply borderline infuriating but I can't pinpoint exactly why that is.

Regardless, here is me, a native speaker, disproving your hypothesis.

I tried the following words in google translate elefantas ailaifantas ailaiphantas elaiphandas elaiphandac.

The suggested detections are ελέφαντας, αιλαιφάντας, αιλαιφάντας, ελαϊφάντας, ελαϊφάντας, however, the translations are elephant, illuminated, illuminated, elephant, elephant respectively. The first is correct. When mapping the roman characters back to greek, there is loss of information, this is seen in the umlaut above iota which makes the pronunciation from ε [e] - like to αϊ [ai̯], and the emphasis denoted via the mark above epsilon (έ).

Notice that all all the words have an edit distance of >=4, a soundex distance of at most 1, and a metaphone distance of at most 1 [1]. The suggested words as I said above are near homophones of the correct word bar a few minor details.

[1] http://www.ripelacunae.net/projects/levenshtein

Re: DALL-E 2 has a secret language

#68
post #34

Shouldn't this be expected to a certain extent? Gibberish has to map _somewhere_ in the models concept space. Whether is maps onto anything we'd recognise as consistent doesn't mean that the AI wouldn't have some concept of where it relates, as other people have noted, the gibberish breaks down when you move it into another context, but who's to say that Dall-E 2 isn't remaining consistent to some concept it understa…

> Shouldn't this be expected to a certain extent?

Not really. It's a stochastic model, so after a bunch of random denoising steps, it could easily just be mapping every bit of gibberish to a random image, and it be vanishingly unlikely for any of them to be similar or the relationship to run in reverse.

Re: DALL-E 2 has a secret language

#69
post #62

Earlier quoted context omitted.

I don't think that's how language detection works, they most likely use the frequencies of n-grams to detect language probability. It's still detected as Greek if you change to "Apoulon vesrreaitais", just because it kind of looks the way Greek words look, not because it resembles any specific word.

You are wrong. Had it been that simple I would __not__ have suggested that and for whatever reason I find your reply borderline infuriating but I can't pinpoint exactly why that is. Regardless, here is me, a native speaker, disproving your hypothesis. I tried the following words in google translate elefantas ailaifantas ailaiphantas elaiphandas elaiphandac. The suggested detections are ελέφαντας, αιλαιφάντας, αιλαιφά…

> for whatever reason I find your reply borderline infuriating but I can't pinpoint exactly why that is.

I guess that says more about you than about my reply. Also, I'm a native speaker as well. That doesn't really have any bearing, my comment above comes from what I know about common implementations of language detection algorithms, not so much from looking at how Google Translate behaves.

Re: DALL-E 2 has a secret language

#70

My first thought upon reading this: what if DALL-E (or a similar AI) uncovers some kind of hidden universal language that is somehow more "optimal" than any existing language? i.e. anything can be completely described in a more succinct manner than any current spoken language. Or maybe some kind of universal language that naturally occurs and any semi-intelligence life can understand it. Fun stuff!

There are a couple sci-fi short stories in the book "Stories of Your Life and Others" by Ted Chiang which explore the idea that highly advanced intelligences might create special languages which accommodate special thoughts which we cannot easily think.
Post reply on HN