Live data from Hacker News

Precursor

blog.cloudflare.com

141–150 of 170 posts

Re: Precursor

#141
post #17

What prevents bots/agents from just adding "jitter" to their movements that mimics how humans move their cursor? I know there are other signals being used but this one in particular seems like it wouldn't be hard to beat with a small amount of sophistication from the bot.

In 2027 how many tokens will we spend to create the jitter, pre-jitter planning, post-jitter verification, and then cloudflare’s inevtiable counter-jitter

"We got this Trace-Buster-Buster-Buster that's gonna bust the Trace-Buster-Buster and bust their .... uh, uh, uh ... Trace!!"

Re: Precursor

#142
post #27
post #17

Earlier quoted context omitted.

In 2027 how many tokens will we spend to create the jitter, pre-jitter planning, post-jitter verification, and then cloudflare’s inevtiable counter-jitter

Someone needs to vibecode a "virtual mouse" tool for the agents to steer instead (semi /s)

Not even. If this is being detected by client-side JS, someone can just reverse-engineer that code, and push a stream of signals into CF to emulate what a human user would generate.

Re: Precursor

#143

Earlier quoted context omitted.

Acting like a human is something scapers already do. Using residential proxies, using latest Chrome user agents, not moving/typing as fast, etc. This is just 1 more layer, moving mouse naturally.

And at some point they'll just use slave labor in some country with lax laws around all that. I'm not sure if I'm talking about the scrapers or Cloudflare at this point. Probably both. Probably the same pool of forced laborers.

A return to the "Mechanical Turk".

I remember using Amazon's Mechanical Turk for just this purpose a decade ago.

With other players in the gig economy really squeezing the workers at the bottom of the system, it could easily get to a point where sitting in front of a computer performing tasks allocated by AI which needs a meat puppet to get around AI-powered Anti-AI might become economically attractive.

Re: Precursor

#144

This (agent detection) is now a kind of emerging space. Obviously it'll get much more important, too. Other products in the space: - Foil ( https://usefoil.com/ ), I'm biased, a friend is building this - Kasada https://www.kasada.io/ - DataDome ( https://datadome.co/ ) - Castle ( https://castle.io/ ) - Fingerprint ( https://fingerprint.com/ ) - HUMAN ( http://humansecurity.com/ ) - Google Cloud Fraud Defense, which i…

happy to offer a counter of some great products for anti-bot defeat:

https://brightdata.com/

https://www.zenrows.com/

https://www.capsolver.com/

https://scrapfly.io/

hundreds of millions of residential ips, human browser fingerprints, custom browser binaries, auto solve of turnstyle, recaptcha v3, kasada, datadome, AWS WAF, etc if they come up.

Re: Precursor

#145

This (agent detection) is now a kind of emerging space. Obviously it'll get much more important, too. Other products in the space: - Foil ( https://usefoil.com/ ), I'm biased, a friend is building this - Kasada https://www.kasada.io/ - DataDome ( https://datadome.co/ ) - Castle ( https://castle.io/ ) - Fingerprint ( https://fingerprint.com/ ) - HUMAN ( http://humansecurity.com/ ) - Google Cloud Fraud Defense, which i…

happy to offer a counter of some great products for anti-bot defeat: https://brightdata.com/ https://www.zenrows.com/ https://www.capsolver.com/ https://scrapfly.io/ hundreds of millions of residential ips, human browser fingerprints, custom browser binaries, auto solve of turnstyle, recaptcha v3, kasada, datadome, AWS WAF, etc if they come up.

Bright Data ranks #1 on Foil's leaderboard [1], but is still detected. ScrapFly is #4 and ZenRows #7. And I guess Capsolver isn't really a scraping thing but is more just for the captcha component.

I think it'd be good if there were more products that did a better job of making an actually undetectable agent, but doesn't seem like any exist yet.

[1]: https://usefoil.com/research/stealth-browser-leaderboard

Re: Precursor

#147

Earlier quoted context omitted.

Gonna zag here. If you take a step back, cloudflare has been paving the path for pay for crawl. I think it's a noble and ambitious goal. While I can understand why you would be alarmed, I can point to almost two decades of lamenting on this forum about how we need better ways of rewarding content creators than ads. Well, this is it. Moreover, these products weren't built in a vacuum. Most threads about Anthropic and…

What's incredible is that you have businesses paying CloudFlare to stop their content being ingested by AIs (OpenAI, Anthropic, Self-operated scrapers). And at the same time, they're paying SEO experts to make that same content easier to be ingested by systems (Google and other Search Engines) which use it for their own AI offerings. Are you going to be able to make your online content available to Google Search but…

Yes. Google obeys robots.txt. Set yours to disallow Google-Extended and they won't train Gemini on it.

Re: Precursor

#148
post #56

Earlier quoted context omitted.

Yes. I am asking exactly that objective question. They are advantageous to leverage in certain situations but essential they are not. We're used to, in the technology industry, looking for or creating problems to solve with services we are aware of. Moving back to necessity and need, do we really? Are we being objective? Most of the time, no.

Theres a few alternatives, but at a minimum yes you probably need their or a competitor's Name Servers and their public DNS. Rolling your own isn't very feasible.

DNS hosting is easy if you did want to roll your own, but everyone one does it, especially your registrar. The more sticky bits of Cloudflare are their cloud services, specifically compute (workers) and database and object storage, all wrapped up into a convenient product (Pages). Also their identity and VPN stuff. They've come a long way since just doing DDoS protection/being a CDN.

Re: Precursor

#149

Earlier quoted context omitted.

happy to offer a counter of some great products for anti-bot defeat: https://brightdata.com/ https://www.zenrows.com/ https://www.capsolver.com/ https://scrapfly.io/ hundreds of millions of residential ips, human browser fingerprints, custom browser binaries, auto solve of turnstyle, recaptcha v3, kasada, datadome, AWS WAF, etc if they come up.

Bright Data ranks #1 on Foil's leaderboard [1], but is still detected. ScrapFly is #4 and ZenRows #7. And I guess Capsolver isn't really a scraping thing but is more just for the captcha component. I think it'd be good if there were more products that did a better job of making an actually undetectable agent, but doesn't seem like any exist yet. [1]: https://usefoil.com/research/stealth-browser-leaderboard

i guess more people should use foil...big recaptcha, cloudflare challenge/turnstyle etc is 95+% bypassed, take your bets on how long this new cloudflare holds out

Re: Precursor

#150
I'm sure they are doing more than looking at mouse movements, but I don't think this is a compelling approach -- I think human movements could be faked pretty easily using a large corpus of real world behavioral data. It's an adversarial game but the level of intelligence we are approaching with AI this can be solved IMO.
Post reply on HN