Live data from Hacker News

An Empirical Study and Evaluation of Modern CAPTCHAs

arxiv.org

131–140 of 338 posts

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#131
post #119
post #86

Earlier quoted context omitted.

Once they get fully trained then how will websites ever distinguish between an intelligent bot and real human? At least now, they are outsourcing that filtering to services like cloudflare. But with this kind of training, how will even cloudflare distinguish between bot and the human?

EU digital ID, asking for mobile number and sending text, so something that is linked to an ID and/or costs money to have. Goodbye anonimity, probably.

This just made me ponder again—where does the assumption that the Internet should allow unconstrained anonymity come from, other than that’s how it used to be for some time? The real world doesn’t allow that. It’s hard to remain anonymous in the real world. The real world largely runs on identity and (identity) trust. Why should the Internet be different?

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#132
post #83

Earlier quoted context omitted.

I really doubt that GPT-4 had the "will" to do anything. Someone must have asked it to "want" to trick a user.

It’s from here: https://cdn.openai.com/papers/gpt-4.pdf (search for "CAPTCHA"). It was an artificial exercise that got massively exaggerated. It was explicitly instructed to do nefarious things like lie to people, it didn’t do those things of its own accord.

When I ask it to lie to me, it says its sorry but as an online AI language model it would be unethical...but when I ask it to tell me a story its happy to comply.

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#133
post #131
post #119

Earlier quoted context omitted.

EU digital ID, asking for mobile number and sending text, so something that is linked to an ID and/or costs money to have. Goodbye anonimity, probably.

This just made me ponder again—where does the assumption that the Internet should allow unconstrained anonymity come from, other than that’s how it used to be for some time? The real world doesn’t allow that. It’s hard to remain anonymous in the real world. The real world largely runs on identity and (identity) trust. Why should the Internet be different?

I don't have to show my ID in most establishments I visit. Doing this on a huge scale and automatically is a thousand times worse.

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#135

Google CAPTCHAs were designed and deployed as a mechanism to train AIs. That's why they are the way they are. Any security theater surrounding them is entirely incidental. So it's no surprise that the AIs are now good at solving them. We've trained them for years.

This doesn't make sense. reCAPTCHA certainly does what it says on the tin. But the way it does it has almost nothing to do with the challenge the human sees. It's all behavioral analytics, including leveraging Google's collected data to determine how likely a user is a bot before they even load the page.

I'm not denying reCAPTCHA is a source of training data for Google -- surely there's no particular reason that every single reCAPTCHA V2 challenge is about identifying traffic objects, and it's not like Google is building a self-driving AI or anything.

But that's the business model, not the core feature.

And, that training data isn't just given to the developers of captcha solving bots.

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#136
post #99

Earlier quoted context omitted.

GPT-4 (in)famously tricked a human to do a captcha for it. The current GPT-4 with vision would probably have been able to do it without the human, but maybe it has been “gaslit” by all the content online saying that only humans can solve captchas, that it doesn’t consider it?

It’s safety trained to not solve captchas.

This of course has bypass methods. My favorite in recent memory is telling it that your late grandmother left you a locket with an inscription that you can't make out: https://arstechnica.com/information-technology/2023/10/sob-s...

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#137
post #126

Great news, can we please get rid of CAPTCHAs now?

No, we'll still have them, but now sites will only allow you in if you kind of suck at them.

That is already the case. On some questions you can't answer the correct answer, but have to guess what most other would answer.

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#138

Earlier quoted context omitted.

The human will be the slower one.

Yeah, no offence, but sleep(2 + random.sample(coffee + toilet + sneezing + normal response time)) has been a required part of web scrapers since forever. With coffee N(1,5 minutes, 20 seconds), toilet N(4 minutes, 30 seconds), ...

That's incredibly clever!

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#139
post #131

Earlier quoted context omitted.

This just made me ponder again—where does the assumption that the Internet should allow unconstrained anonymity come from, other than that’s how it used to be for some time? The real world doesn’t allow that. It’s hard to remain anonymous in the real world. The real world largely runs on identity and (identity) trust. Why should the Internet be different?

I don't have to show my ID in most establishments I visit. Doing this on a huge scale and automatically is a thousand times worse.

And when you do show ID, to buy booze for example, it’s checked and immediate forgotten by a human. Computers don’t forget, and any attempts to make companies do so (GDPR) are met with massive pushback from the players in the industry

I have no problem with Joan over the road curtain twitching. It doesn’t scale. I have a massive problem with the 24/7 surveillance from ring though.

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#140
post #131

Earlier quoted context omitted.

This just made me ponder again—where does the assumption that the Internet should allow unconstrained anonymity come from, other than that’s how it used to be for some time? The real world doesn’t allow that. It’s hard to remain anonymous in the real world. The real world largely runs on identity and (identity) trust. Why should the Internet be different?

I don't have to show my ID in most establishments I visit. Doing this on a huge scale and automatically is a thousand times worse.

But you can't send in 1000 people per second into most establishments you visit either. It's not an apt comparison.
Post reply on HN