Live data from Hacker News

Why can’t a bot tick the 'I'm not a robot' box?

quora.com

191–200 of 647 posts

Re: Why can’t a bot tick the 'I'm not a robot' box?

#191
post #6

Earlier quoted context omitted.

Those images are infuriating! Click all boxes with traffic lights. Ok, well, this one box just barely contains the bottom right corner of the traffic light. Click. Nope, that little corner didn't count. Try again. Ok, well on this one, the right side of the traffic light is only barely over the line, so I won't click it. Nope, that sliver of the light mattered this time. MF!

I assumed the infuriating ambiguity is intentional, in order to train some algorithm they need to know what the prevailing human correct judgement is in dicey situations

I don't think it's intentional -- it probably just emerges from the training process.

I'm guessing they do something like load up a batch of images and once N people agree on one, record the answer and remove it from the rotation. You end up left with the ambiguous images where people couldn't agree.

Re: Why can’t a bot tick the 'I'm not a robot' box?

#192
post #149
post #80

I feel a sense of dread whenever I see this box. Is it going to let me through, or am I going to spend the next few minutes futilely clicking signs and lights, only to give up and leave the site?

I'm willing to bet if you record how long you spend clicking signs and lights on average, it's going to be more like several seconds than a few minutes. This must be hyperbole. There are the outlier cases where it gets fairly annoying but otherwise, I'm not sure I understand the hostility toward such a benign system that actually does thwart bots very effectively.

> I'm willing to bet if you record how long you spend clicking signs and lights on average, it's going to be more like several seconds than a few minutes.

How much, and at what odds?

The length of the average Google CATPCHA has been steadily going up for me. I haven't pulled out a stopwatch, but I do count how many image sets I go through. I basically never get through on the checkbox unless I've done one on another site shortly before. I sometimes succeed after one set, but if I don't it's consistently 3+. The worst case I've seen was 10 layers of slow-loading images without success, at which point I gave up and tried on another device. (If it had been a site I didn't need, I'd have given up after 5 - which I do fairly often, so I don't have an average count needed to succeed!)

I'm fully aware that average users don't have this much trouble, or people would be furious. But I also see that captcha ramps up to an extremely long process in the face of even modest privacy-protection efforts like not running Javascript or allowing third party trackers by default. (God forbid you're using a VPN for any reason.) It's not assessing your humanity but your familiarity, using the same fingerprinting tools as any site that wants to track you.

Spam is a real problem, and a hard one to solve, but I admit I'm hostile to Google's captcha. Partly because it really is a significant time sink for me. Partly because it lacks any progress indicator or fallback option so it's an indefinite hurdle to accessing sites I'm already committed to using. But largely because, despite what I really believe are good intentions, it's yet another force pushing people to give up privacy and even security if they want websites to work tolerably.

Re: Why can’t a bot tick the 'I'm not a robot' box?

#193
post #7

Earlier quoted context omitted.

This might surprise you, but it actually has to do with what traffic coming out of TOR looks like. Well in excess of 90% of traffic coming out of TOR is spam, bots, malicious, or some combination! Google isn't going out of their way to punish you for trying to protect your privacy. They're trying to stop unwanted traffic. By unfortunate happenstance, you appear to be disguising yourself in the exact same way a shocki…

This. The reality is, Google (and Cloudflare, and everyone else trying to block scrapers and malicious traffic) use heuristics that boil down to "99% of our traffic behaves like this". If you go out of your way to fall into the 1%, e.g. using Tor, disabling Javascript, randomizing your user-agent, etc., you're going to get CAPTCHAed.

Yeah, blending in seems to work better in many cases. Remember the guy who sent a bomb threat over TOR? The only reason he was caught so quickly was because he's the only person on the organisation's network to have accessed TOR before the incident.

Re: Why can’t a bot tick the 'I'm not a robot' box?

#194

Earlier quoted context omitted.

I'm convinced the ambiguity is intentional. What I don't get is what answer they expect in those scenarios.

I always figure they're looking for a population consensus. They're doing image recognition at scale and these are clearly ambiguous, hard images to classify. They could easily have a few people at Google say, "I determine this is a storefront" and make that the "correct" answer, but I think they're more interested in a consensus of what most "normal" people would classify as a storefront, especially in potentially-v…

What they're actually getting though is the population consensus of what normal people believes Google's image classifier believes. The system incentivizes users to reinforce misconceptions their classifier has.

Does this look like a mountain to you? https://0x0.st/zzvr.jpg

Google's image classifier would think that's a mountain. If you disagree, google will classify you as a robot. After failing these sort of challenges a few times the user decides to play along and tell google what they think google wants to hear, rather than the truth.

Re: Why can’t a bot tick the 'I'm not a robot' box?

#195

I'm not convinced that picking the pictures has anything to do with actually convincing google if you're a bot or not. I mean sure, it's an indicator. But I _know_ that I can pick the right pictures of school buses and store fronts every freaking time, so that's only a very small indicator. More likely, the majority of the algorithm is devoted to the "fingerprint" of your browser. If you have adblock running, you may…

My theory is that the challenges are actually just to slow down humans trying to do repetitive sensitive tasks (sign ups for instance).

Re: Why can’t a bot tick the 'I'm not a robot' box?

#197

As a user who's constantly clicking on the crosswalk or storefront images you can't help but to think that you're essentially working for free training Google's machine learning models by providing them with supervised data points.

I've been thinking about this a lot lately. Where is our compensation? It's our time and brain power training Google's AI that will one day be sold back to us. I'm really not into this.

Because Google can extract value from captchas, it makes world-class captchas and bot detection AI available to every webmaster for free. I don't know what that level of service would otherwise cost, but it almost certainly wouldn't be affordable for low-traffic blogs and the like, which would end up vulnerable using weaker captchas or trying to roll their own. Everywhere else the cost would just get passed on to users.

I don't love the compromise of paying for things with my data or by training Google's AI, but it's hard to say users aren't getting anything out of it. That said, I do miss the old reCaptcha.

Re: Why can’t a bot tick the 'I'm not a robot' box?

#199
post #173

Earlier quoted context omitted.

I've been thinking about this a lot lately. Where is our compensation? It's our time and brain power training Google's AI that will one day be sold back to us. I'm really not into this.

So you want compensation because your data is used along with millions of others to train an algorithm to distinguish if a bot or a real human to provide a service to you? Nice.

Yes if the data is of value. They don't give this data out publicly. Open source the data or pay.

Re: Why can’t a bot tick the 'I'm not a robot' box?

#200

As a user who's constantly clicking on the crosswalk or storefront images you can't help but to think that you're essentially working for free training Google's machine learning models by providing them with supervised data points.

That's what ReCaptcha always was... it was originally a known and another unsure text blurb from scanned books/text documents. Now it's street signs etc.

It's actually pretty interesting to see how the captchas have evolved as, presumably, Google decides "OK, we have enough data to consistently identify this thing" and moves on to the next challenge.

I recall the modern (non-text) captchas used to be cars pretty much every time. Then, the images started getting grainier as they apparently wanted to improve their recognition in different conditions. Then crosswalks and store fronts became quite common, eventually with the same kinds of noise distorting images. Now I've started seeing things like buses, bridges, motorcycles, bicycles, etc. It feels like they've finished getting enough data for improving Google Maps and have begun moving towards collecting data for their self-driving car projects.

Post reply on HN