Live data from Hacker News

Shirt Without Stripes

github.com

411–420 of 639 posts

Re: Shirt Without Stripes

#411
post #197

This problem is known as "attribution" - you have a "no" or "without" in the sentence, but you don't know where it belongs. One could (and one does) argue that the problem cannot be solved with statistical methods (ML), especially not in any domain where accuracy is required, such as medical recored analysis: "no evidence of cancer" and "evidence of no cancer" are very different things. Zooming out, the language fiel…

The # 1 google result for “Shirt without stripes” is this very own HN post https://www.google.com/search?q=%22Shirt+without+Stripes%22&

Google is in the business of producing quick results to sell ads while keeping the cost low. If anyone does better, they will likely do it using a costlier algorithm, and if there is more margin in reselling tangible products than ads then Amazon is incentivized to use costlier algorithms that are more accurate. But I think the margin in both cases does not justify the use of algorithms that may be more accurate in some small percentage of cases but a lot more costly.

Re: Shirt Without Stripes

#412

Earlier quoted context omitted.

> The model just needs to be able to understand the important meaning of "no" in here, in the context of the whole sentence It's easy to say, isn't it? Unfortunately, sticking the word "just" in there doesn't affect the difficulty. I do it all the time, too. That said, "meaning" is not statistical.

Why is it hard? Honestly. Add a "t" to "no" and you get a basic Boolean query.

Boolean queries aren't statistically analyzed; they're parsed and algebraically evaluate. So your point seems to support the opposite thing you're supporting.

Re: Shirt Without Stripes

#413

Earlier quoted context omitted.

You may be right. It's been bugging me since I posted earlier on so I fired up a VPN with an endpoint in Japan, along with a private browsing session in Firefox, to see if I got different results. As it happens the results were interesting : - If I entered "person" I'd see a mix of images substantially similar to what I saw using google.co.uk up to and including Terry Crews, which was frankly a little weird, and othe…

This seems like you're attributing motive to google here, but I don't believe that's right. For example, Terry Crews appears in the query "person" because his "TIME Person of the Year 2017 Interview" article was very popular online. I get a lot of Greta Thunberg because she was TIME Person of the Year 2019 and received similar online attention because of Donald Trump. The TL;DR of it is that google crawls the interne…

> Google doesn't "know" what you're trying to search. It's a giant pattern matching game that slices and dices and rearranges text to find the closest match.

I'm not disputing that, and it certainly explains why it's "good enough" for somes search queries whilst being totally gimpy for others.

My understanding was that Google does prioritise what it's classified as local search results though, on the basis that they're likely to be more relevant.

Re: Shirt Without Stripes

#414

Why should it not be possible to solve this with statistical methods? The model just needs to be able to understand the important meaning of "no" in here, in the context of the whole sentence. I would guess that most modern NNs from the NLP area (Transformer or LSTM) would be able to correctly differentiate the meaning. The problem is, I think there is no fancy NN (yet) behind Google search, and the other web searche…

"there is no evidence of cancer" and "there is evidence of no cancer" are two different statements with different meaning, so it's more complex a task than just understanding the importance of "no" in a sentence. It's involves semantic analysis of the sentence. The paper I linked to below describes a technique they call "deep parsing." Check it out for more context.

Lexical analysis is surely clever enough to make "no evidence" and "no cancer" the atoms used and should even differentiate "no cancer" from "no, cancer" pretty easily? Are we really still at 80s chatbot level functionality when it comes to string parsing?

Related aside: It frustrates me no end that spellcheck still doesn't appear to use any probablistic considerations, like Markov chains, to determine the intended word. And that when I click the next to last letter to make an adjustment it doesn't then change the suggestions to alternate endings, etc.. Perhaps newer devices than I have do this.

Re: Shirt Without Stripes

#415
post #331

Earlier quoted context omitted.

If that's the case that's a big problem, because human children are trivially capable of both formulating and understanding uncommon sentences which still make sense. It might be hard to come up with examples on the spot, but in everyday life you will routinely come across things you need to refer to by negation which are relatively uncommon.

Human children are also capable of walking, but it’s one of the hardest problems in robotics. Things brains can do intuitively, they can do because they have millions of years of evolution behind them.

Of the things the human brain does, I wouldn't say walking is very interesting.

The most amazing property in my opinion is the fact that it trains itself. (Whereas neural networks are trained by external systems).

I suspect it's related to imprinting. To take the example of filial imprinting, the brain must have some hardcoded notion of what a parent looks like. Then this is used to to build a parent detector, and the hardcoded notion is thereafter discarded. Then the newly learnt parent detector is used in reinforcement learning (near parent = good). Keep in mind that this all happens just after birth or hatching, before the visual centres of the brain have had any chance to train.

Really cool stuff.

Re: Shirt Without Stripes

#416

Earlier quoted context omitted.

I couldn't quite believe your comment when I read it so I did a Google image search for "person" and the results weren't a lot better than you'd suggested. Mostly white men, a few white women, a very few black women, a handful of Asians, and multiple instances of Terry Crews. The net result of that Google search, combined with the "Shirt Without Stripes" repo, leaves me even more unimpressed with the capabilities of…

If you really want to be disappointed, search for [doctor] and [nurse]. Unless things have really changed, [doctor] will be mostly white men and [nurse] will be mostly white and Filipino women. But don't blame the AI. The AI has no morality. It simply reflects and amplifies the morality of the data it was given. And in this case the data is the entirety of human knowledge that Google knows about. So really you can't…

> The AI has no morality. It simply reflects and amplifies the morality of the data it was given.

Key point right there. Unless Google is deliberately injecting racial and/or gender bias into their code, which seems extremely far fetched (to put it kindly), the real fault lies with us humans and what we choose to publish on the web.

Re: Shirt Without Stripes

#417
post #197

This problem is known as "attribution" - you have a "no" or "without" in the sentence, but you don't know where it belongs. One could (and one does) argue that the problem cannot be solved with statistical methods (ML), especially not in any domain where accuracy is required, such as medical recored analysis: "no evidence of cancer" and "evidence of no cancer" are very different things. Zooming out, the language fiel…

Yes, you have points, but they break down here: “Shirt -stripes” is unambiguous to a system, yet the first result on Amazon(.ca) is a striped shirt, and the 3rd is sweatpants.

As someone else said you don't know that's wrong for what Amazon are optimising. If they find people [with your background profile] who buy shirts are susceptible to buying sweatpants, they might also find that if the seed you with "sweatpants" as an idea up front that the repeated presentation of sweatpants in "people who bought X also bought Y" sections is more effective.

That's the sort of thing I'd expect Amazon to be doing?

Re: Shirt Without Stripes

#418

Earlier quoted context omitted.

I couldn't quite believe your comment when I read it so I did a Google image search for "person" and the results weren't a lot better than you'd suggested. Mostly white men, a few white women, a very few black women, a handful of Asians, and multiple instances of Terry Crews. The net result of that Google search, combined with the "Shirt Without Stripes" repo, leaves me even more unimpressed with the capabilities of…

If you really want to be disappointed, search for [doctor] and [nurse]. Unless things have really changed, [doctor] will be mostly white men and [nurse] will be mostly white and Filipino women. But don't blame the AI. The AI has no morality. It simply reflects and amplifies the morality of the data it was given. And in this case the data is the entirety of human knowledge that Google knows about. So really you can't…

What bias? Who is biased? Quick duckduckgoing indicates there are far more male than female doctors in the US. So statistically, it would be correct to return mostly male doctors in an image search. If you want a photo of a specifically gendered doctor, it's not hard to specify. Not really seeing a problem here.

Re: Shirt Without Stripes

#420
Fun experiment on Google:

- Shirt Without Stripes: shirts where the description contains both "without" and "stripes". Example: a shirt without collar, with stripes.

- "Shirt Without Stripes": a mess, with and without stripes, suggesting an unusual search query. In fact, the linked article site is the first result in web search.

- Stripeless shirt: sexy women in strapless shirts

- "stripeless shirt": pictures of Invader Zim...

- "stripeless" shirt: mostly shirts without stripes, but there are some shirts with stripes that are described as stripeless...

The last one may give us a hint at the problem. If you have to mention a shirt is without stipes, you are probably comparing is to a shirt with stripes. For example imagine a forum, some guy is posting a picture of a shirt with stripes, I can expect some people to ask questions like "do they sell this shirt without stripes"? Or maybe the seller himself may have a something like "shirt without stripes available here (link)" in the description. So the search engines tie "shirt without stripes" to pictures of shirts with stripes.

I remember an incident where searching for "jew" on Google led to antisemitic websites. The reason was simply that that exact word was rarely used in other contexts. Mainstream and Jewish source tend to use the words "jews" and "jewish" but not "jew". And because Google doesn't look at the dictionary meanings of words but rather what people use them for, you get issues like that.

Post reply on HN