Live data from Hacker News

The Lone Banana Problem in AI

digital-science.com

61–70 of 109 posts

Re: The Lone Banana Problem in AI

#62

"Three cats in a trenchcoat" is a good one. Three cats in a trenchcoat standing on each other's shoulders, pretending to be a human, Vincent Adultman style. You can get them standing side by side wearing trenchcoats. You can get a cat pyramid or a cat totem or a stack or tower of cats (though often the ability to count to three is then lost). You can't get them to share the trenchcoat. Nothing like that occurred in a…

It's mostly the generation engine, not the logical AI. Here's what GPT-4 says when asked to describe. > As an AI, I can't actually see the painting, but based on the title and the dimensions you've provided, I can imagine a possible description. > The painting, titled "Three cats in a trenchcoat standing on each other's shoulders, pretending to be a human, Vincent Adultman style," is a large piece, standing 6 feet ta…

Not sure why this post is getting down voted? It's spot on - the limitation lies in relatively small text encoders within stable diffusion/mid journey. Larger language models contribute to better image generation (see: DeepFloyd). Google's CapPa also recently showed an alternative to the contrastive learning of CLIP at any param count.

Until larger models are released, there are "hacks" like this available: https://bair.berkeley.edu/blog/2023/05/23/lmd/ - GPT4-generated bounding boxes guiding stable diffusion

Re: The Lone Banana Problem in AI

#63

I wish this article was just 3 paragraphs. The verbose writing style was a little tiring, I found myself scrolling impatiently to find what the actual "Lone Banana Problem" was.

Maybe fittingly, it is a writing style I have associated with AI (ChatGPT in particular). Here, the article doesn't have the ChatGPT "signature", but it has the same tendency of having a lot of grammatically correct filler.

Re: The Lone Banana Problem in AI

#64
Whatever the critiques of length, this is a good write up on 'bias', that is not about race or politics, and thus a good neutral example about how training sets can skew results in un-intended ways.

I didn't think it was that long, it might be that some are cherry picking the points they find interesting and think the rest could have been edited out.

I could have stood even more discussion on how this 'blind spot' in the AI model is very much akin to our own blind spots.

Re: The Lone Banana Problem in AI

#65
post #51

Earlier quoted context omitted.

I read almost the whole article thinking there might be something more in there. There was mention of TCAV but that was about it.

I am mystified. The page is not long, the text is not complex. The message is obvious and it is contained in the title as shown here on HN. The first pic and its caption make it plain. The next group of 4 pics and their SINGLE SENTENCE caption spell it out clearly. As for "almost the whole article" -- it is short! It took about 90sec to read the whole page, top to bottom. "AI" chat bots can't count. This is well know…

> As for "almost the whole article" -- it is short! It took about 90sec to read the whole page, top to bottom.

Cool. The article has 2348 words, so that's around 26 words per second, or ~1560 words per minute.

According to some speed reading pages I randomly found via Google the consensus seems to be that 1000+ words per minute is quite exceptional. https://irisreading.com/what-is-the-average-reading-speed/ puts average adult reading speed at around 300 WPM, which you casually exceed by 400%.

Far be it from me to suggest that your estimate is inaccurate, but at least know that not everybody reads at that speed :-)

Re: The Lone Banana Problem in AI

#66

I wish this article was just 3 paragraphs. The verbose writing style was a little tiring, I found myself scrolling impatiently to find what the actual "Lone Banana Problem" was.

This is the biggest communication problem most people have. Say it plain and simple. No one wants to read your train of thought.

I used to write emails to my boss as summary, then clearly mark what follows as the (boring) details. I got a lot of promotions that way.

Re: The Lone Banana Problem in AI

#67
post #51

Earlier quoted context omitted.

I read almost the whole article thinking there might be something more in there. There was mention of TCAV but that was about it.

I am mystified. The page is not long, the text is not complex. The message is obvious and it is contained in the title as shown here on HN. The first pic and its caption make it plain. The next group of 4 pics and their SINGLE SENTENCE caption spell it out clearly. As for "almost the whole article" -- it is short! It took about 90sec to read the whole page, top to bottom. "AI" chat bots can't count. This is well know…

My point exactly. This article could have been written: "AI is biased against drawing just one banana" and I would have preferred it. And I said "almost" because I skimmed two of the paragraphs in the middle that seemed unlikely to contain information.

Re: The Lone Banana Problem in AI

#68

I wish this article was just 3 paragraphs. The verbose writing style was a little tiring, I found myself scrolling impatiently to find what the actual "Lone Banana Problem" was.

Agree. I was a couple of pages in before I learned what the "Lone Banana Problem" is. That took two paras, and was followed by a lot of sophomoric philosophical noodling.

[dead]

Re: The Lone Banana Problem in AI

#69
I remember about 10 years ago in the context of using neural network for image classification, they had a study of what the AI actually saw that justified its decision. This, by the way, is how we got "Deep Dream".

For "dumbbell", for the AI, no dumbbell was complete without a muscular arm holding it. That's because in most images in the training dataset have the dumbbells being held, so the system integrated the arm in the pattern. I guess it is the same idea here, most pictures of bananas show several of them, so for the AI, bananas are things that don't go alone.

Re: The Lone Banana Problem in AI

#70
post #51

Earlier quoted context omitted.

I am mystified. The page is not long, the text is not complex. The message is obvious and it is contained in the title as shown here on HN. The first pic and its caption make it plain. The next group of 4 pics and their SINGLE SENTENCE caption spell it out clearly. As for "almost the whole article" -- it is short! It took about 90sec to read the whole page, top to bottom. "AI" chat bots can't count. This is well know…

My point exactly. This article could have been written: "AI is biased against drawing just one banana" and I would have preferred it. And I said "almost" because I skimmed two of the paragraphs in the middle that seemed unlikely to contain information.

But that isn't the point.

The point is, LLMs look smart but they are not, and an easily-verifiable data point is that they can't count.

Post reply on HN