Live data from Hacker News

hCaptcha now runs on fifteen percent of the internet

hcaptcha.com

221–230 of 380 posts

Re: hCaptcha now runs on fifteen percent of the internet

#222

Earlier quoted context omitted.

How'd you do it?

I'm not OP but there are cheap solving services that will fill them in for you. The cost is trivial and they have decent APIs for automatic integration into scrapers.

many captchas let you fill them in once and then use the resulting cookie to do whatever else you like on the site. Just manually copy the cookie to your bot hosts.

Re: hCaptcha now runs on fifteen percent of the internet

#223

Earlier quoted context omitted.

I built an alternative[0] that takes a proof of work approach. As a site owner you set the difficulty that makes sense for you: so perhaps you would want 20 seconds of computation before you can submit. The nice thing is that this can happen entirely in the background while the user fills in the form. Also with multiple requests from the same IP in a short timespan, the difficulty increases. There are downsides to to…

First one to make this mine an altcoin for proof-of-work wins. But seriously I like the idea, although it seems trivial for someone to attack a protected site by exhausting its subscription level? Are there any protections against that?

We don't disable the service if a protected site goes over their limit.

Right now we manually look at the limits and are reasonable with overages - also we can see how many captchas were unsolved.

Re: hCaptcha now runs on fifteen percent of the internet

#224

Earlier quoted context omitted.

I built an alternative[0] that takes a proof of work approach. As a site owner you set the difficulty that makes sense for you: so perhaps you would want 20 seconds of computation before you can submit. The nice thing is that this can happen entirely in the background while the user fills in the form. Also with multiple requests from the same IP in a short timespan, the difficulty increases. There are downsides to to…

So your solution is to technically waste electricity to replace captcha? It's for sure an interesting concept, the first point and low-end devices requiring 20+ seconds to pass are not a very good points to sell your service.

You're right that there is an electricity cost to solving this type of captcha - the same as there is an electricity cost to loading 2MB of JS+images and clicking the pictures with the fire hydrants (and the infrastructure behind that). It's hard to estimate how they compare (and what value you assign to the human labor performed and privacy loss).

20 seconds would be a fairly high difficulty. It's up to the site owner to decide what makes sense for them.

If anybody comes up with a useful computational task with a small bundle size that can be verified cheaply that would be the holy grail - until then the computation is only there as a form of hashcash.

Re: hCaptcha now runs on fifteen percent of the internet

#225
post #96

I see the product is offered in 2 plans, though the only paid plan is enterprise without publicly disclosed pricing. Anyone has info whether it's affordable also for small and bootstrapped businesses or it's primarily focusing on larger enterprises?

Last time I enquired they quoted starting $999/mo for 10m verifications...

Thank you. That's not a small sum, though they probably have some internal tiers per number of requests so for lower traffic scenarios the monthly price might be more affordable.

Re: hCaptcha now runs on fifteen percent of the internet

#226
post #126
post #48

I dislike the widespread use of captcha regardless of provider. I realize anything connected to the internet will be subject to automated abuse, and it's impossible to run some types of services without taking some steps to defend against it, but it seems to me there's usually a way to handle that without invading the user's privacy or wasting their time. The exact details will vary based on the type of service, of c…

> it seems to me there's usually a way to handle that without invading the user's privacy or wasting their time As much as I agree with your dislike of captchas, I don't think this is true at scale (unless universal online identities existed, which could and should include anonymous identifiers by design). When you need to accept information from anonymous users (comments, votes, forms, registrations), there's no way…

CAPTCHA does not scale. CAPTCHA spams real people with requests and wastes my VALUABLE time, and still labels disabled people as subhuman. It's offensive. It's ineffective. It's outdated.

It's reaching a point where encapsulating a VPN with anti-captcha is something I'd pay for.

Re: hCaptcha now runs on fifteen percent of the internet

#227
post #28
post #4

Worth noting that this title is primarily due to Cloudflare having switched to them from ReCAPTCHA, and Cloudflare is... well, relatively popular, to say the least. I'm curious what kind of data may exist on the experience of switching for larger providers; do the users like it? how much more/less time do they spend solving? do they care, let alone even notice that it's not Google's ReCAPTCHA? Regardless, as ReCAPTCH…

Disclaimer: I've been an engineer at hCaptcha for a few years now building out the service. I'm just as interested in you as hearing about customer and user success/pain stories! > Worth noting that this title is primarily due to Cloudflare having switched to them from ReCAPTCHA, and Cloudflare is... well, relatively popular, to say the least. That's definitely a part of it, but we also have a number of other large s…

I usually just bounce when I see a captcha (if I get one, I usually get a string of them, so I don’t bother).

However, I checked secondary markets where you can pay a human to solve a captcha.

It takes a professional captcha solver 70 seconds to solve an hCaptcha but only 15-20 seconds to solve a reCaptcha. Is that typical? That seems horrible.

The market rate for a captcha solution is 1-3 cents, which is clearly worth it, until you think of the ethics of paying someone slave wages so you can browse the internet slowly, but at least without breaking concentration.

Have you considered a more ethical approach, like micropayments that go to charity or something?

Re: hCaptcha now runs on fifteen percent of the internet

#228

Earlier quoted context omitted.

I built an alternative[0] that takes a proof of work approach. As a site owner you set the difficulty that makes sense for you: so perhaps you would want 20 seconds of computation before you can submit. The nice thing is that this can happen entirely in the background while the user fills in the form. Also with multiple requests from the same IP in a short timespan, the difficulty increases. There are downsides to to…

So your solution is to technically waste electricity to replace captcha? It's for sure an interesting concept, the first point and low-end devices requiring 20+ seconds to pass are not a very good points to sell your service.

Right? I think they’re describing those crypto mining scripts people were being inflicted with a while back :)

Re: hCaptcha now runs on fifteen percent of the internet

#229
post #48

I dislike the widespread use of captcha regardless of provider. I realize anything connected to the internet will be subject to automated abuse, and it's impossible to run some types of services without taking some steps to defend against it, but it seems to me there's usually a way to handle that without invading the user's privacy or wasting their time. The exact details will vary based on the type of service, of c…

Here's a thought experiment. This one requires some long-term thinking, outside the box and well past recent history and the status quo.

What if the majority internet usage is non-interactive, from so-called "bots", what we may refer to as "automated use". Google and Facebook, among others, rely on the use of automation and "bots". The non-interactive clients ("bots") being used by these companies are not asked to solve captchas. (In turn, after collecting data from public sources, these websites attempt to prohibit the use of automation by their users wishing to access it. What is interesting is that neither company provides any definition of "automated" nor any clearly stated limits on the speed at which a user may access resources or the quantity of resources they may access in a stated time period. One might be apt to find such limits associated with an "API".)

In 2013 an Incapsula report suggested that the majority of internet usage is in fact automated and not "malicious"^1 -- what if public information sources on the internet catered to the use of automation rather than trying to limit such use, e.g., with speed bumps^2 like "captchas". What if servers treated all clients equally, instead of having data forcibly collected by a few large clients that receive preferential treatment, then siloed and protected from "automation". What effects would this have on "centralisation" and levelling the playing field.

"Do not ask for permission, ask for forgiveness." What does it really mean when applied to the internet. Perhaps it means there is an endemic lack of clarity about "the rules". Prohibiting "automation" is far too vague and in many cases it makes no sense. The growth of computers and the internet is the growth of automation. Both servers and clients may have concerns about resource utilisation. Websites do not ask for permission when they decide to use large amounts of the user's computer resources.

Consider that a Google could not exist without being "given permission" to use automation. Does the GoogleBot have to solve captchas. No automation means no company such as this could exist. How useful would the web be without anyone being able to use automation to create an index. Based on the HN comments about web search I have read over the years, I would guess that for many commenters, it means the usefulness of the web would be dramatically reduced.

Imagine an automation-friendly internet. The truth is, I think (the data shows) we already have one, except we are in denial that "the rules" actually allow it. An early metaphor for internet and web use was "surfing". It may be that those who are constantly fighting against automation are fighting against the waves instead of riding them. Time will tell. It stands to reason, IMO, that every internet user, whether a server or a client, should be expected to use automation.

1. https://www.incapsula.com/blog/bot-traffic-report-2013.html

2. An early metaphor for the internet was a "superhighway". Speed bumps would seem out of place on a superhighway.

Re: hCaptcha now runs on fifteen percent of the internet

#230

As someone who scrapes, captcha's are pretty silly. One of the sites we scrape implemented hCaptcha, and it was a breeze to get around. There are a few things that make my life more difficult, but captchas aren't one of them, and nothing can stop scraping altogether.

Meh, there's always going to be a longtail of targeted abuse so it's not much to boast over. Xrumer software in 2001 could even let you sit at your computer and fill out those common PHP-lib captchas (like on EZBoard) while Xrumer spammed internet forums and blogs. You could even hire a cubicle farm of humans to manually abuse a web service. Captchas filter out the 90% bulk of automated abuse. Btw, web scraping is on…

That makes sense that there's longtail abuse, thanks. What would you say is more harmful abuse? Spamming endpoints for SQL and other injection attacks?
Post reply on HN