Live data from Hacker News

Dazed and Confused: A Large-Scale Real-World User Study of ReCAPTCHAv (2023)

arxiv.org

31–40 of 65 posts

Re: Dazed and Confused: A Large-Scale Real-World User Study of ReCAPTCHAv (2023)

#32
> reCAPTCHAv2 checkbox presents itself as a complete vulnerability disguised as a security tool ... It can be concluded that the true purpose of reCAPTCHAv2 is as a tracking cookie farm for advertising profit masquerading as a security service.

Woof, some super high quality "research" going on here. Their basis for this claim is that they found a paper from 2016 where the authors did some cookie aging, found it worked and passed that on to Google, who then fixed it. The authors cite this "blatant vulnerability" but then seem to assume that ReCAPTCHA being versioned means that v2 has never changed since it launched? Then they discover the amazing fact that CAPTCHAs require energy and disappear down a conspiracy theory rabbit hole before concluding the entire product - that people pay money for - is actually useless.

The reality is that academics can't contribute much directly to modern anti-spam work and that has been true for a long time. They aren't genuinely spammers so tend to act in ways that don't trigger anti-spam systems (which is correct behavior). The possibility that their simulations of bad activity aren't accurate enough just doesn't occur to them, so then they run around reporting "vulnerabilities". Usually it's easier to tweak the system to make them happy than have them go to the press and cause a scene, so they think they're doing useful work, but it's actually just distracting employees from the actual job of spam fighting. Back in 2020 Twitter publicly lost patience with this type of university output and slammed it as "extremely limited" and "behind the curve" [1]. Citing an eight year old paper as if it's still relevant is a classic example of this problem. The publish-and-cite model is designed for investigating natural laws, not black box commercial systems that can change on a weekly basis or faster. Their supervisor should really have stopped them citing something that old.

The other problem is that because researchers don't understand spammers very well they tend to assume that anything they can do, spammers can also do. In reality spam is purely a profit/loss optimization problem. Spammers are business people who want to make money. The goal is to throw just enough speedbumps in their path to make their profit margin negative, at which point they give up and the spam goes away, without also tipping your own numbers into the red by over-spending or hurting the UX too badly. Making things tougher than necessary risks pushing your own margins negative, making things easier than necessary risks users getting annoyed at the spam and leaving. The goal is to ride the edge of the wave as closely as possible. When spammers find a trick that works they keep it as a trade secret and sell the results of using the trick, not the trick itself, and so disrupting the relationships between black market suppliers is an important tactic. CAPTCHAs in particular are just a throttle intended to slow down low grade spam; if you're actually spending 20-30 seconds automatically solving image puzzles with an 80% success rate you haven't beaten anything, because slowing you down that much was the goal in the first place.

This environment is nothing like academia, where people are getting funded by the government to spend months doing work basically for the heck of it. In academia it's reasonable to spend a year gathering data, running experiments, training ML models, solve a tiny subset of the actual problem spammers face, ignore the cost of doing all that because it's government subsidized, and then announce you've defeated something. If they were really spammers they'd given up and moved on to a different scheme long before that (in a surprising number of cases, this new scheme will be getting a proper job). So whether they realize it or not, all this line of work really does is funnel tax money into the spam ecosystem - exactly what we don't want.

Source: worked in web anti-spam for a while.

[1] https://blog.x.com/en_us/topics/company/2020/bot-or-not "the threat has evolved and the narrative on what’s actually going on is increasingly behind the curve."

Re: Dazed and Confused: A Large-Scale Real-World User Study of ReCAPTCHAv (2023)

#33
post #15

Earlier quoted context omitted.

It takes a script kiddy considerably more effort to circumvent a captcha than just automating a site via curl or chromium. This difference is the increase in cost of an attack. This is the security gain.

It's now so trivial to solve them, and extremely cheap. You can install a chrome extension, give some solving service a $1 and basically never need to reload.

Still makes a ddos much more difficult

Re: Dazed and Confused: A Large-Scale Real-World User Study of ReCAPTCHAv (2023)

#34
post #14

Earlier quoted context omitted.

Captcha is nothing like a lock. It's a little guy that gives you a run around before you get to insert your key. It does very little to stop the bad actors (if there's a payday at the other end of the runaround, they'll do it), but annoys (and is a slap in the face for) every single legitimate user.

Captcha is like locks on soap shelves. It's friction on people who want to buy soap and can be trivially defeated if you want to steal the soap.

Not only that, but adding insult to injury - the whole captcha process is being abused by Google for their surveillance capitalism.

Just load a captcha-blocked page in a fresh browser with no google accounts every being used - you got like a 1% chance of getting through. Login to google, and BOOM you're in.

Re: Dazed and Confused: A Large-Scale Real-World User Study of ReCAPTCHAv (2023)

#35

Earlier quoted context omitted.

Captcha is like locks on soap shelves. It's friction on people who want to buy soap and can be trivially defeated if you want to steal the soap.

Not only that, but adding insult to injury - the whole captcha process is being abused by Google for their surveillance capitalism. Just load a captcha-blocked page in a fresh browser with no google accounts every being used - you got like a 1% chance of getting through. Login to google, and BOOM you're in.

From the study:

"In terms of cost, we estimate that – during over 13 years of its deployment – 819 million hours of human time has been spent on reCAPTCHA, which corresponds to at least $6.1 billion USD in wages. Traffic resulting from reCAPTCHA consumed 134 Petabytes of bandwidth, which translates into about 7.5 million kWhs of energy, corresponding to 7.5 million pounds of CO2. In addition, Google has potentially profited $888 billion USD from cookies and $8.75-32.3 billion USD per each sale of their total labeled data set."

Re: Dazed and Confused: A Large-Scale Real-World User Study of ReCAPTCHAv (2023)

#36
post #31

I use AI to solve my captchas now. Pay for the service from nopecha, the internet has been killed for VPNs without it

The failure of cryptocurrency to provide a micropayments platform for the web is quite frustrating. That $5 per month should be flowing to content producers, but it's paying the energy bill of some GPU farm instead.

Re: Dazed and Confused: A Large-Scale Real-World User Study of ReCAPTCHAv (2023)

#37
post #11

I abandon 1 in 10 captchas recently. They're becoming so abstract and confusing, few are on sites that I critically need to persevere with. Of the ones I do, I need to regenerate about a third as I'm slightly colour blind, but apparently enough to not see whatever they need me to see. Mostly I won't sign up to a service that has a captcha unless I know in advance it's a service I need (unlikely), it's an immediate cl…

[deleted]

Re: Dazed and Confused: A Large-Scale Real-World User Study of ReCAPTCHAv (2023)

#40

To play devil's advocate; it is better to have some kind of fallback when the automated, non-interactive verification does fail. Anecdotally from many posters here and as I've personally experienced, automated verification (as attempted by some captchas which entirely eliminate any user interaction) is not foolproof, and an escape hatch is necessary from time to time. While it could be argued that this leaves the cap…

[deleted]
Post reply on HN