Doesn't ignoring headers and taking someones site down fall afoul of the CFAA? Especially given how comments here are showing that this is a recurring issue? Depending on which side of the fence you're on, there could be standing to go after FB for this either for money, or for their lawyers to help set a precedent to limit CFAA further.
When someone does this to Facebook its malicious and they go to jail. WHen Facebook does it to someone else... "oops".
A Facebook crawler was making 7M requests per day to my stupid website
81–90 of 416 posts
Re: A Facebook crawler was making 7M requests per day to my stupid website
#82Re: A Facebook crawler was making 7M requests per day to my stupid website
#83This is not how the mail works, where someone needs to buy a stamp to spam you.
Re: A Facebook crawler was making 7M requests per day to my stupid website
#84Doesn't ignoring headers and taking someones site down fall afoul of the CFAA? Especially given how comments here are showing that this is a recurring issue? Depending on which side of the fence you're on, there could be standing to go after FB for this either for money, or for their lawyers to help set a precedent to limit CFAA further.
When someone does this to Facebook its malicious and they go to jail. WHen Facebook does it to someone else... "oops".
Re: A Facebook crawler was making 7M requests per day to my stupid website
#85Re: A Facebook crawler was making 7M requests per day to my stupid website
#86Earlier quoted context omitted.
This pattern needs to die. Whenever I paste a link in any sort of messaging app - iMessages to Slack - it puts a thumbnail and summary, polluting the entire conversation with tons of noise. Messaging apps have no business in looking up the URL. Just let it pass as a link. Fuck everything about this and we need to push back on this nonsense.
Part of the goal here is to prevent people from clicking on malicious links. A preview helps with that. Or, more simply, you can't be rickrolled if you know the destination in advance.
Re: A Facebook crawler was making 7M requests per day to my stupid website
#87Earlier quoted context omitted.
I disagree completely, I find it extremely useful. Most links are unreadable (e.g. an ID to a cloud file), and are truncated anyways even if readable. The preview gives me the title of the webpage, which is infinitely more useful -- especially letting me know whether it's a document I've already seen (and don't need to open) or something new, and if so, what.
To me it pollutes the discussion away from the messages and puts a bunch of thumbnails where it could just be a clean text conversation. The thing that bothers me is also that I, as a user, have no choice - it just does it automatically. Do you like IRC?
Re: A Facebook crawler was making 7M requests per day to my stupid website
#88Make sure you file a bug - there are a myriad of sources internally, but we can often hunt it down easily enough (assuming it gets triaged to eng). Important info is the host of the urls being crawled, and the User Agent attached to the requests (headers are also good too). A timeseries graph of the hits (with date & timezone specified) can also help.
Re: A Facebook crawler was making 7M requests per day to my stupid website
#89We've had the same issue. They were doing huge bursts of tens of thousands of requests in very short time several times a day. The bots didn't identify as FB (used "spoofed" UAs) but were all coming from FB owned netblocks. I've contacted FB about it, but they couldn't figure out why this was happening and didn't solve the problem. I found out that there is an option in the FB Catalog manager that lets FB auto-remove…
Thanks man! I'll have a look.
I just sent you an e-mail, you can also reply to that instead if you prefer not to share those details here. :-)
Re: A Facebook crawler was making 7M requests per day to my stupid website
#90Make sure you file a bug - there are a myriad of sources internally, but we can often hunt it down easily enough (assuming it gets triaged to eng). Important info is the host of the urls being crawled, and the User Agent attached to the requests (headers are also good too). A timeseries graph of the hits (with date & timezone specified) can also help.
Fuck this culture of "it's up to the victim of our fuckup to file a bug report with us".