Live data from Hacker News

Why can’t a bot tick the 'I'm not a robot' box?

quora.com

531–540 of 647 posts

Re: Why can’t a bot tick the 'I'm not a robot' box?

#531

I think recaptcha and captcha in general are very overused today, as they cause way too much inconvenience for the user. Why discriminate against robots so much? What's so wrong about crawling or using automated tools? With today's networks and hardware performance most websites shouldn't concern themselves with denial-of-service type of attacks, unless they're past a certain threshold of popularity.

At work we have a scraper that likes to use a particularly expensive search query using hundreds of different ips dozens of times per second, all so that they can scrape data that is freely available from us as an XML feed at a different url. Every time I found a way to fingerprint and block them, they'd change their bot to avoid detection. Captcha for rate limiting seemed like the least bad option. However, eventual…

Why not add a "PLEASE USE THE XML DUMMY!" line to the returned results when you figure out it's them?

Re: Why can’t a bot tick the 'I'm not a robot' box?

#532
post #526
post #86

Earlier quoted context omitted.

I think quora over states what Google looks at by a wide margin, just try to access a captcha in incognito, they won't have access to as much info as they do on you and yet you're still presented with the same level of captcha (if not more of them, which is to be expected)

Sometimes just checking the checkbox is enough. Sometimes you need to identify cars and store fronts. I think the better Google knows who you are, the more likely just the checkbox is going to be enough. If you go incognito, you have to train their neural nets, if you give up your privacy, you get in for free. The clever part from Google's perspective is that you have to trade one of these things to Google in order t…

There are many services out there that can solve Google's recaptcha for fraction's of a penny. When someone puts one up, they can make things more expensive, and perhaps sometimes uneconomical, but in general, the cost is low (~$2.00 for 1,000 recaptchas).

When someone uses a recaptcha, they should think about why they are doing so. It's one thing to use it to save a business model, but it's another to use it to protect information that should be free anyway. The elephant in the room is government data. Many government agencies think that selling their data can be a nice source of side revenue, and a recaptcha is a good way of enforcing it. In reality, they just increase the costs for everyone, and those with means can obtain the information while those without means cannot.

Governments need to release their data, freely, without captchas or fees for single users and bulk users, no exceptions.

Re: Why can’t a bot tick the 'I'm not a robot' box?

#534

The box has made browsing using TOR insufferable! It fusses and makes me click storefronts and traffic lights until I run out of patience and close out of whatever webpage I was trying to visit. I assume it has to do with a lack of Google cookies on the browser, essentially punishing me for trying to protect my privacy.

> punishing me for trying to protect my privacy. TOR doesn't protect your privacy, it just lumps you in with—and makes you indistinguishable from—the worst crap on the internet. If you don't want to be treated like crap, don't try to blend in with the crap.

This is frankly an idiotic statement. Many people use Tor out of principle, not because of their desire to "blend in with the crap".

Re: Why can’t a bot tick the 'I'm not a robot' box?

#535
post #66

Earlier quoted context omitted.

I don't think anyone is concluding that.

I think a lot of people come to exactly that conclusion.

Why would they? Most people cognizant of these terms knows a bot generates more traffic than a human; that’s the point of most bots.

Re: Why can’t a bot tick the 'I'm not a robot' box?

#536

Earlier quoted context omitted.

This is very hostile to people who use screen readers.

I've found that the tech industry often is. Trying to get managers to set aside time to iron out accessibility issues is like pulling teeth. Trying to get other developers to take it seriously is almost as bad. Often you count yourself lucky if the bare legal minimum is implemented. Accessibility is very important, and if accessibility features are implemented well they'll often be useful even to people without disab…

> Can you even imagine 21st century university architecture department that didn't cover ADA compliance? That'd be unthinkable.

I can easily imagine it: architecture departments from universities in other countries don't necessarily have to cover compliance with USA laws.

Re: Why can’t a bot tick the 'I'm not a robot' box?

#537

As a user who's constantly clicking on the crosswalk or storefront images you can't help but to think that you're essentially working for free training Google's machine learning models by providing them with supervised data points.

I've recently made a discovery that pleases my petty side.

You know how they usually give you several questions to solve, even if you're quite convinced you solved a question correctly?

Turns out if you click randomly, they keep showing you new questions as well. If, after a handful of purposely wrong answers, you answer one correctly, they let you through.

I now purposely mess up the answers a few times. It seems neither slower nor faster than actually taking the time to do it right, but it takes less mental load, and it makes me not feel like doing slave labour for a machine.

Re: Why can’t a bot tick the 'I'm not a robot' box?

#538

Earlier quoted context omitted.

It's actually pretty interesting to see how the captchas have evolved as, presumably, Google decides "OK, we have enough data to consistently identify this thing" and moves on to the next challenge. I recall the modern (non-text) captchas used to be cars pretty much every time. Then, the images started getting grainier as they apparently wanted to improve their recognition in different conditions. Then crosswalks and…

Those really grainy images, they always seemed like they had to have noise added on purpose? It was like really badly processed film grain, but if the images were enlargements then wouldn't they be pixelated? Stuff you get now often requires cultural information, like "sidewalk" isn't a cross-cultural name, I'd guess almost everyone knows it, but meh. What classes as a store, is a lawyers office a store? Also, I seem…

Yeah, a lot of the storefront ones are just plan hit things randomly until it lets you through - how am I supposed to know if a building with some writing on it in Korean is a store or something else? Are you supposed to include the poles in traffic lights or not?

The V2 was just annoyingly badly designed because the questions were badly put.

Re: Why can’t a bot tick the 'I'm not a robot' box?

#540
post #202

Earlier quoted context omitted.

That's what ReCaptcha always was... it was originally a known and another unsure text blurb from scanned books/text documents. Now it's street signs etc.

they stopped using the text because some forum campaign that promoted typing cursewords instead of the unkown word. they probably started showing cursewords on the rendered search highligths on google books. there was a decent write up from a whitehat showing the damage, but I can't find it

> they stopped using the text because some forum campaign that promoted typing cursewords instead of the unkown word. they probably started showing cursewords on the rendered search highligths on google books.

If I remember correctly, Google later on also sometimes showed two "known" words or, if they had actual other evidence that you are human, two unknown words.

Post reply on HN