Live data from Hacker News

An Empirical Study and Evaluation of Modern CAPTCHAs

arxiv.org

331–338 of 338 posts

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#331
post #323

Earlier quoted context omitted.

Am I the only one paranoid enough to think that this means Apple is now indexing even the text content of images stored on it's users computers?

That's literally a feature of the platform. If you open the photos app and type text in it will give you the photos containing that text. If your concern is "apple is harvesting my data" then no. All of apple's various analysis systems ("AI") are entirely local. This does mean you get a bunch of duplicated work as every device redoes the same analysis but on the other hand it saves you from "how do we defend against…

> All of apple's various analysis systems ("AI") are entirely local.

Even if that were true (I couldn't say, and I don't think anyone who doesn't have access to Apple source code and production systems could either) that wouldn't preclude Apple harvesting the results of said AI analysis.

In fact, doing the analysis on users' devices would represent a shift of that processing from cloud to edge, representing a significant savings for Apple or anyone else in a similar position.

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#332

Earlier quoted context omitted.

A simple zero-knowledge credential system isn't sufficient. It would need to embed some kind of protections to limit how often it could be used, to detect usage of the same credential from multiple (implausibly far apart) IP addresses. There would need to be extremely sophisticated reputation scoring and blocklisting to quickly catch people who built fake identities or stole them. And even with every one of those pro…

Yes, I wonder how feasible it is to do that while still protecting state of being anonymous. And what if you develop this very sophisticated system of reputation score, what if bad actors find a way to still perfectly abuse it, e.g. they pay for desperate people for the IDs and then stay just within the limits ever so slightly. Would you be able to easily iterate on the system when that happens to make it more secure…

There are a lot of risks here and I think it’s very challenging to build something anonymous that can deal with (say) Google’s current level of fraudulent behavior, let alone what we’re likely to see in the future.

Regarding the IP address question, I’d assume you could decouple the IP address verification portions from the “know who the person is” portions with some clever multi-party computation. Someone always has to know your IP address, but it doesn’t have to be the same person you’re talking to. (Think of Tor as an inspiration here.)

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#333
post #323

Earlier quoted context omitted.

That's literally a feature of the platform. If you open the photos app and type text in it will give you the photos containing that text. If your concern is "apple is harvesting my data" then no. All of apple's various analysis systems ("AI") are entirely local. This does mean you get a bunch of duplicated work as every device redoes the same analysis but on the other hand it saves you from "how do we defend against…

> All of apple's various analysis systems ("AI") are entirely local. Even if that were true (I couldn't say, and I don't think anyone who doesn't have access to Apple source code and production systems could either) that wouldn't preclude Apple harvesting the results of said AI analysis. In fact, doing the analysis on users' devices would represent a shift of that processing from cloud to edge, representing a signifi…

Apple's literal marketing message is that they do all the processing on device. It's not a shift for Apple, as apple has never done this analysis in the past, and only started doing the analysis once it could do it locally.

The use cases we're talking about also don't work on a cloud based analysis, as you can't have text selection block on network uploads (generally slower than downloads), and it would require uploading every image you open to apple which would presumably be a lot of traffic, and an obvious privacy nightmare. It would also break for users who turn on the e2ee everything mode for iCloud.

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#334
post #333

Earlier quoted context omitted.

> All of apple's various analysis systems ("AI") are entirely local. Even if that were true (I couldn't say, and I don't think anyone who doesn't have access to Apple source code and production systems could either) that wouldn't preclude Apple harvesting the results of said AI analysis. In fact, doing the analysis on users' devices would represent a shift of that processing from cloud to edge, representing a signifi…

Apple's literal marketing message is that they do all the processing on device. It's not a shift for Apple, as apple has never done this analysis in the past, and only started doing the analysis once it could do it locally. The use cases we're talking about also don't work on a cloud based analysis, as you can't have text selection block on network uploads (generally slower than downloads), and it would require uploa…

It's wonderful (for Apple's bottom line) that you believe these things. I assume you can show source code and provide access to production systems to verify?

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#335
post #319
post #63

Earlier quoted context omitted.

> hasn’t yet been able to crack vision driven self driving But they have? For years Google Street view has read signs, house numbers, phone numbers of businesses, etc. from the environment. It is safe to assume they have this built into Waymo as well. I assume you might be trying to reference "vision only" self-driving, which is a fantasy made up by Elon Musk because nobody would sell him LiDAR sensors cheaply. https…

This is a meme. “Sour grape Elon, touting vision because no one will sell him LiDAR sensors. Which are the gold standard sensors that solve self driving.” How exactly does LiDAR tell you whether the thing in question can move (dog) or not (trash can)? How does it allow a neural net to infer intent? You’ll actually have to solve vision. Even if you had LiDAR. There’s no way around it. And once you’ve solved it, LiDAR…

> How exactly does LiDAR tell you whether the thing in question can move

LIDAR is continuously scanning, usually multiple times a second. It is irrelevant if the object is a trash can or a dog if it has a trajectory into the street.

> And once you’ve solved it, LiDAR becomes superfluous.

Only if you have the low Musk level standards of simply being equal to a human. There are plenty of jobs robots can do better than humans and driving is one of them. But it does require LIDAR and/or radar.

https://abc7news.com/tesla-s-autopilot-self-driving-car-offi...

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#336

Earlier quoted context omitted.

Wow, that never happened to me and is unacceptable. Was that for the infotainment only or the drive train? Just for others, they are separate systems, you can even safely reboot the infotainment (main display with maps, music etc) if you need to while driving, as it doesn't affect the drive train. I'm guessing it was not the drive train which would be incredibly dangerous.

Yeah, it didn't affect the drive train, and it was also quite quick - less than a minute between when the screen went dark and when it had finished rebooting and sent notifications that an update had been installed. So presumably just an infotainment update as you said; I didn't try to dig into exactly what the update included though.

It will also reboot the infotainment sometimes (but not always) when it crashes.

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#337

Earlier quoted context omitted.

In the us, I noticed that grocery stores increasingly scan your drivers license (my state has bar codes). I think it's probably a way to keep clerks from passing someone through who is not quite 21 (a different captcha!). I have wondered if they keep the scan or does the state? I asked and the random hourly worker there said they don't.

Do those grocery stores still scan your drivers license (or I guess any other ID) if you don't buy alcohol?

no, they only scan if you buy booze.

Re: An Empirical Study and Evaluation of Modern CAPTCHAs

#338
post #212

Earlier quoted context omitted.

mCaptcha is interesting, but I wonder what its energy impact would be on a sufficiently large deployment, e.g imagine we replaced all reCAPTCHAs with mCaptcha.

Author of mCaptcha here o/ mCaptcha uses PoW and that is energy inefficient, but it not as bad as the PoWs used in blockchains. The PoW difficulty factor in mCaptcha is significantly lower than blockchains, where several miners will have to pool their resources to solve a single challenge. In mCaptcha, it takes anywhere between 200ms to 5s to solve a challenge. Which is probably comparable to the energy used to train…

> Which is probably comparable to the energy used to train AI models used in reCAPTCHA.

I had not considered that. Naturally, we're just speculating here, but yeah that does sound plausible.

I was also no aware of the "hard" 5s bound (which you seem to have tested on a normal smartphone setup); sounds neat.

Post reply on HN