Live data from Hacker News

The Lone Banana Problem in AI

digital-science.com

81–90 of 109 posts

Re: The Lone Banana Problem in AI

#81
I unsubscribed from Midjourney because I couldn't direct it to place an object in a specific place.

It's impossible to do something like, "A sphere on the left of the picture." and have it understand it.

After having iterative conversations in ChatGPT, Midjourney felt like it has an extremely poor grasp of language.

Re: The Lone Banana Problem in AI

#82

> AIs, at their current level of development, don’t perceive objects in the way that we do – they understand commonly occurring patterns. You see this claim everywhere - that AI operates on statistics and patterns and not actual understanding. But human understanding is entirely about statistics and patterns. When a human sees a collection of particles and recognizes it as, say, a car, all they are doing is recognizi…

After a decade of billions invested in systems with models trained for object recognition in traffic, these models still struggle with object permanence. Why should we expect some rando painter models to do better? They'll paint a lamppost through a car and not notice anything wrong with it.

The day where these models can show shoes under a fence and a reflection of the person behind the fence in an opposing shop window? That day will come. But not on the current crop.

Re: The Lone Banana Problem in AI

#83

Earlier quoted context omitted.

Thought process is fine, but lead with the plain and simple version. Offer context after you've introduced the topic and described what you're talking about.

I think when you're trying to be convincing of a counter intuitive or controversial point, it helps to start with less objectionable parts of the argument first. If an argument involves several steps, A to B to C ending in D, but D is wildly counter intuitive, some audiences, hearing D first will reduce their likelihood of actually integrating the consequences of A and so on. Maybe someone who's done more research on…

I'm not so sure. A shining example is the essay "One Self: The Logic of Experience"[0]. The first paragraph is just 3 sentences, and the first sentence lays it out:

"I believe that we are all the same person. How could anybody come to think that? In this paper I shall try to explain how it happened to me." [https://www.researchgate.net/publication/233329805_One_Self_...]

Re: The Lone Banana Problem in AI

#84
post #71
post #65

Earlier quoted context omitted.

> As for "almost the whole article" -- it is short! It took about 90sec to read the whole page, top to bottom. Cool. The article has 2348 words, so that's around 26 words per second, or ~1560 words per minute. According to some speed reading pages I randomly found via Google the consensus seems to be that 1000+ words per minute is quite exceptional. https://irisreading.com/what-is-the-average-reading-speed/ puts aver…

Yeah, I am a speedreader, and I did used to be able to do circa 3000wpm at a push. Good to know that even as a myopic 55YO I can still do half that without trying. Points of comparison: I originally read Hal Clement's classic Mission of Gravity in about 25min, and Joseph Conrad's Victory in about 3 hours.

Right, so your 90 seconds is actually closer to 10 minutes for most people. That's a lot of time to communicate "AI tends to render one banana as two bananas", which was OP's point.

Yes, there were pictures that communicated the same thing, but people naively thought that those 2348 words may actually hold some additional information, as otherwise ca 2300 of them would be completely superfluous. But they were wrong; again, OP's point. Not sure what's so mystifying by this.

Re: The Lone Banana Problem in AI

#85

I wish this article was just 3 paragraphs. The verbose writing style was a little tiring, I found myself scrolling impatiently to find what the actual "Lone Banana Problem" was.

This is the biggest communication problem most people have. Say it plain and simple. No one wants to read your train of thought.

Apparently those kinds of people consider you to be the one with the communication disorder. I consider them incapable of writing an abstract.

Re: The Lone Banana Problem in AI

#86

I wish this article was just 3 paragraphs. The verbose writing style was a little tiring, I found myself scrolling impatiently to find what the actual "Lone Banana Problem" was.

Do you really not know how to skim effectively? Took me just a few seconds.

Re: The Lone Banana Problem in AI

#87

"Three cats in a trenchcoat" is a good one. Three cats in a trenchcoat standing on each other's shoulders, pretending to be a human, Vincent Adultman style. You can get them standing side by side wearing trenchcoats. You can get a cat pyramid or a cat totem or a stack or tower of cats (though often the ability to count to three is then lost). You can't get them to share the trenchcoat. Nothing like that occurred in a…

While I agree with you, and the article, that the larger problem is that the AI model simply hasn't experienced enough data to get an accurate grasp on the situation, or that the data was labelled in a way that influences the model's understanding, I think the problem here may be a human one. In the article update, the author says that they managed to craft a prompt that got the result they want by specifying the banana must be on its own. The model knows what a banana on its own looks like, but the human is expecting the model to "do what I mean" and getting frustrated when the model "did what I said". Now, I'll admit that I skimmed the last third of the article, but I didn't see any mention that things like Stable Diffusion and Midjourny have a syntax, saying "a single banana casting a shadow on a grey background" is different from "((single banana)), casting shadow, hard light, dramatic, grey background" for example.

Re: The Lone Banana Problem in AI

#88
post #16

Earlier quoted context omitted.

On /r/midjourney subreddit there are posts like "the most stereotypical person in [state/country]", "what midjourney thinks professors look like based on their department". You can see some interesting biases in the training data. https://www.reddit.com/r/midjourney/top/?sort=top&t=year

I wouldn't use "bias" here. It's a very ambiguous term in machine learning.

Not sure what a better term would be... Just being more specific might help, like "bias in the training data"?

Re: The Lone Banana Problem in AI

#89

Midjourney solution: "--no bunch, two, multiple": https://i.imgur.com/QnJmLr0.jpeg These models have a tendency to move towards the average, especially if unprompted. As we see here, sometimes even if prompted otherwise. They just have to get better or need a more precise interface like the —no :). We also couldn't have "a man crawling" before and now we can: https://i.imgur.com/ycVpk3i.jpeg

I agree. While the article is correct that this problem of subtle bias exists, the solution for most cases is to realise that the computer "does what you say, not what you mean".

Re: The Lone Banana Problem in AI

#90
post #84
post #71

Earlier quoted context omitted.

Yeah, I am a speedreader, and I did used to be able to do circa 3000wpm at a push. Good to know that even as a myopic 55YO I can still do half that without trying. Points of comparison: I originally read Hal Clement's classic Mission of Gravity in about 25min, and Joseph Conrad's Victory in about 3 hours.

Right, so your 90 seconds is actually closer to 10 minutes for most people. That's a lot of time to communicate "AI tends to render one banana as two bananas", which was OP's point. Yes, there were pictures that communicated the same thing, but people naively thought that those 2348 words may actually hold some additional information, as otherwise ca 2300 of them would be completely superfluous. But they were wrong;…

I don't think that the time difference is that big. I'm fast, but I don't think I'm that much faster than an average skim reader.

As I said, the title of the article and a couple of image captions are enough to get the gist, and as such, I find it baffling that so many people seem to have totally failed to understand it.

Post reply on HN