Finally, a competent response that doesn't leave the users hang out to dry.
I'm so tired of seeing incompetents with measures like "blackhole 2 continents" deployed even outside active attacks.
31–40 of 73 posts
Finally, a competent response that doesn't leave the users hang out to dry.
I'm so tired of seeing incompetents with measures like "blackhole 2 continents" deployed even outside active attacks.
Earlier quoted context omitted.
I've never seen the phrase "manual transmission user-agent". Using your own browser yourself is the new stick shift. Love it.
It feels like everyone's rebuilding their own desktop experience. Kind of Minecraft with folders and text files. The really interesting part of this is how little people talk about what they're doing, and it doesn't feel secretive in any way.
Earlier quoted context omitted.
Just curious, if you're tolerant of scraping, do you make an archive of all your content available so that scraping is unnecessary, and if so do the scrapers prefer that?
Current evidence is that scrapers mostly aren't nearly considerate or sophisticated enough to take an "archive of all content" option if one exists. See https://people.kernel.org/monsieuricon/creepy-crawlies which describes how the https://git.kernel.org gets hammered by crawlers all the time even though you could run a single `git clone` and get the data that way instead.
For somebody who knows a bit how things are set up, or is willing to spend 10 minutes researching, it's a no-brainer that you can just "git clone" entire linux kernel development history, or download entire wikipedia [0].
Alas, large number of scrapers are not willing to spend those 10 minutes, it would appear. So, here we are.
Earlier quoted context omitted.
Just curious, if you're tolerant of scraping, do you make an archive of all your content available so that scraping is unnecessary, and if so do the scrapers prefer that?
Current evidence is that scrapers mostly aren't nearly considerate or sophisticated enough to take an "archive of all content" option if one exists. See https://people.kernel.org/monsieuricon/creepy-crawlies which describes how the https://git.kernel.org gets hammered by crawlers all the time even though you could run a single `git clone` and get the data that way instead.
Earlier quoted context omitted.
Current evidence is that scrapers mostly aren't nearly considerate or sophisticated enough to take an "archive of all content" option if one exists. See https://people.kernel.org/monsieuricon/creepy-crawlies which describes how the https://git.kernel.org gets hammered by crawlers all the time even though you could run a single `git clone` and get the data that way instead.
This is exactly the problem, unfortunately. For somebody who knows a bit how things are set up, or is willing to spend 10 minutes researching, it's a no-brainer that you can just "git clone" entire linux kernel development history, or download entire wikipedia [0]. Alas, large number of scrapers are not willing to spend those 10 minutes, it would appear. So, here we are. [0] https://dumps.wikimedia.org/
Earlier quoted context omitted.
That has only partially mitigated much smaller attacks (residential proxy scraping etc) on my employer's site. We're currently on the "Business" plan, but I'm coming to the conclusion that we need to upgrade to the "Enterprise Advantage" plan for the JA3/4 fingerprinting and detection ID features. I get put off by "Contact Sales" pricing.
Here's my take: * JA3s are mostly useless. JA4s supersede them entirely. * Using JA4s in rate limits is pretty useful and helps a lot against proxy scraping. It was not very helpful in this attack. * Bot detections are somewhat helpful but they don't solve scrapers/attacks by themselves. They're useful as a 2nd/3rd data point (eg. low bot score + bot detection + something else)
Earlier quoted context omitted.
Here's my take: * JA3s are mostly useless. JA4s supersede them entirely. * Using JA4s in rate limits is pretty useful and helps a lot against proxy scraping. It was not very helpful in this attack. * Bot detections are somewhat helpful but they don't solve scrapers/attacks by themselves. They're useful as a 2nd/3rd data point (eg. low bot score + bot detection + something else)
isn't JA4 also useless because it is so easy to spoof tls. For eg. cycletls for nodejs etc..
First, find out who's on the other end of a few hundred IP addresses. Start with ones in the US. Sue for damages. Use discovery to find out what's on the other end. Sue the maker of that device. If it turns out to be an appliance or smart TV, it may be possible to consolidate cases into one case against the manufacturer. Criminal negligence, tort interference with contract, harassment, Computer Fraud and Abuse act violation... Maybe a restraining order prohibiting the sale of "smart TV" known to be able to host attacks. Have imports seized by Customs and Border Protection. That would get a manufacturer's attention.
The manufacturer's EULA will not help the manufacturer, because the plaintiff, the party being attacked, is not a party to the EULA at all.
Yippee, free load test!