Live data from Hacker News

I’m not a human: Breaking the Google reCAPTCHA [pdf]

blackhat.com

51–60 of 70 posts

Re: I’m not a human: Breaking the Google reCAPTCHA [pdf]

#51
post #48

Earlier quoted context omitted.

I'm pretty sure this is the advertised purpose of it. It stops bots AND helps machine learning.

*AND helps a company make profit by training their proprietary models. If they were open, it would be a lot better (considering all the training is done by volunteers, too)

I like to misidentify things to give them bad data.

Re: I’m not a human: Breaking the Google reCAPTCHA [pdf]

#52
post #31

"We ran our captcha-breaking system against 2,235 captchas, and obtained a 70.78% accuracy" That's more impressive than it sounds. I'm pretty sure 70.78% is more accurate than I am with reCAPTCHA manually. A lot of the captcha's presented are very fuzzy, or have ambiguous questions, etc.

Indeed. It's good that these guys are white-hats because they could have made a killing selling to spammers (for as long as they could go undetected, which could've been a while).

Going rate for humans breaking captchas is like $1/1000 captchas solved. A fully automated service could make some money, sure, but not exactly a killing.

Re: I’m not a human: Breaking the Google reCAPTCHA [pdf]

#53
post #40
post #36

Earlier quoted context omitted.

The little captcha box does some processing before deciding what to show you. If you look suspicious, it can show you a more complex challenge. If you look like a normal browser, it might show you a house number to read (to improve its Maps product maybe). If you already have a cookie set because you already proved you're human, maybe it's just that check-box that says "I'm a human".

Unless you're able to explicitly state how this is done and/or how to trick it in think you're "high risk" - then seems like speculation; yes, I'm aware Google's said this, but never seen an proof of it. I've found bug in the past in the system that were easy to fix, told Google, but the bugs never were fixed.

It's not speculation. It's section 2 of the paper, titled "Analyzing Risk Analysis System."

Re: I’m not a human: Breaking the Google reCAPTCHA [pdf]

#54
post #48

Earlier quoted context omitted.

I'm pretty sure this is the advertised purpose of it. It stops bots AND helps machine learning.

*AND helps a company make profit by training their proprietary models. If they were open, it would be a lot better (considering all the training is done by volunteers, too)

uhh. Everything they do is arguably for profit. That's why Google is an LLC. What do you expect them to do, put an "* this is done to earn us money" disclaimer on everything?

Re: I’m not a human: Breaking the Google reCAPTCHA [pdf]

#55
post #51
post #48

Earlier quoted context omitted.

*AND helps a company make profit by training their proprietary models. If they were open, it would be a lot better (considering all the training is done by volunteers, too)

I like to misidentify things to give them bad data.

Well, I do the same. I won’t destroy any of their dataset, or even have a measurable impact – but in turn, I won’t have a measurable impact in positive direction either.

If they want to make a profit from me, they should pay me, or they can license correct CAPTCHA results from me under GPL.

Re: I’m not a human: Breaking the Google reCAPTCHA [pdf]

#56
post #52

Earlier quoted context omitted.

Indeed. It's good that these guys are white-hats because they could have made a killing selling to spammers (for as long as they could go undetected, which could've been a while).

Going rate for humans breaking captchas is like $1/1000 captchas solved. A fully automated service could make some money, sure, but not exactly a killing .

Yeah; that's the problem...you can make, with tons of effort, a fairly decent system to tell robots and humans apart, but it's much much more difficult to tell humans trying to do the thing directly from those solving the challenge remotely. It's an arms race of economics; the challenge has to be difficult enough that it slows down humans in sweatshops to the point where it makes the whole enterprise not worthwhile for the abusers while not pissing off your actual users. It also has to resist automated malicious use. Quite a tall order.

The best I've seen are the "which of these photos show mountains"-type challenges. I'd imagine that solving 5 rounds of those would take too long to make it worthwhile for spammers, but I'd also imagine lots of legitimate users getting irked at going through that to fill out a form.

Re: I’m not a human: Breaking the Google reCAPTCHA [pdf]

#57
post #55
post #51

Earlier quoted context omitted.

I like to misidentify things to give them bad data.

Well, I do the same. I won’t destroy any of their dataset, or even have a measurable impact – but in turn, I won’t have a measurable impact in positive direction either. If they want to make a profit from me, they should pay me, or they can license correct CAPTCHA results from me under GPL.

This argument doesn't really make sense to me. What browser, OS and device are you using to post this comment? Someone made profit when you bought the device, bought the OS and/or browser.

A healthy system needs some kind of motivation. In economies, that is profits/money. What's wrong with that? (I know I am being simplistic here but...)

Re: I’m not a human: Breaking the Google reCAPTCHA [pdf]

#58
post #40
post #36

Earlier quoted context omitted.

The little captcha box does some processing before deciding what to show you. If you look suspicious, it can show you a more complex challenge. If you look like a normal browser, it might show you a house number to read (to improve its Maps product maybe). If you already have a cookie set because you already proved you're human, maybe it's just that check-box that says "I'm a human".

Unless you're able to explicitly state how this is done and/or how to trick it in think you're "high risk" - then seems like speculation; yes, I'm aware Google's said this, but never seen an proof of it. I've found bug in the past in the system that were easy to fix, told Google, but the bugs never were fixed.

I've experienced this first-hand. I used to always get the checkbox, but then one day I had to download a large series of related files from a website that used reCAPTCHA. I did all the clicking manually, but in a very repetitive and bot-like fashion, by opening a series of 10 new tabs at a time and then performing the same series of clicks on each tab to get to the download link. After a few minutes, I stopped getting checkboxes and started getting increasingly more difficult CAPTCHA challenges.

Re: I’m not a human: Breaking the Google reCAPTCHA [pdf]

#59

I am the only one who thinks google reCAPTCHA is just a tool that Google uses to train Machine Learning Algorithms? First it was used to help Google learn how to Read, now its learning to detect object, landscapes, ...

I'm pretty sure this is the advertised purpose of it. It stops bots AND helps machine learning.

The advertised purpose was to digitize the worlds books and help libraries, digital humanities and the world.

This ended, and all trace of public good was erased (Check archive.org if you dont believe me) when Google execs mandated that only things that made money can be supported.

Re: I’m not a human: Breaking the Google reCAPTCHA [pdf]

#60
post #31

"We ran our captcha-breaking system against 2,235 captchas, and obtained a 70.78% accuracy" That's more impressive than it sounds. I'm pretty sure 70.78% is more accurate than I am with reCAPTCHA manually. A lot of the captcha's presented are very fuzzy, or have ambiguous questions, etc.

>I'm pretty sure 70.78% is more accurate than I am with reCAPTCHA manually.

exactly. Many reCAPTCHA are beyond simple recognition and make me start guessing. I expect we'll see new type of reCAPTCHA - you're a human if you make mistake and robot if correct answer is typed in :) Similar to those 1x1 images not visible to humans, yet visible to the robots.

Post reply on HN