A Facebook crawler was making 7M requests per day to my stupid website
61–70 of 416 posts
Re: A Facebook crawler was making 7M requests per day to my stupid website
#62Earlier quoted context omitted.
I disagree completely, I find it extremely useful. Most links are unreadable (e.g. an ID to a cloud file), and are truncated anyways even if readable. The preview gives me the title of the webpage, which is infinitely more useful -- especially letting me know whether it's a document I've already seen (and don't need to open) or something new, and if so, what.
To me it pollutes the discussion away from the messages and puts a bunch of thumbnails where it could just be a clean text conversation. The thing that bothers me is also that I, as a user, have no choice - it just does it automatically. Do you like IRC?
Re: A Facebook crawler was making 7M requests per day to my stupid website
#63Re: A Facebook crawler was making 7M requests per day to my stupid website
#64Earlier quoted context omitted.
This pattern needs to die. Whenever I paste a link in any sort of messaging app - iMessages to Slack - it puts a thumbnail and summary, polluting the entire conversation with tons of noise. Messaging apps have no business in looking up the URL. Just let it pass as a link. Fuck everything about this and we need to push back on this nonsense.
A funny story (well, funny because it didn't happen to me) was chronicled on a recent episode of Corey Quinn's "Whiteboard Confessions" -- https://www.lastweekinaws.com/podcast/aws-morning-brief/whit... -- where Slackbot's auto URL unfurling feature "clicked" an SNS alert unsubscribe link when a tech posted the a report including that link into a slack channel. If nobody lost their job over this triggering a SEV-1 al…
Re: A Facebook crawler was making 7M requests per day to my stupid website
#65We've had the same issue. They were doing huge bursts of tens of thousands of requests in very short time several times a day. The bots didn't identify as FB (used "spoofed" UAs) but were all coming from FB owned netblocks. I've contacted FB about it, but they couldn't figure out why this was happening and didn't solve the problem. I found out that there is an option in the FB Catalog manager that lets FB auto-remove…
Re: A Facebook crawler was making 7M requests per day to my stupid website
#66Earlier quoted context omitted.
This pattern needs to die. Whenever I paste a link in any sort of messaging app - iMessages to Slack - it puts a thumbnail and summary, polluting the entire conversation with tons of noise. Messaging apps have no business in looking up the URL. Just let it pass as a link. Fuck everything about this and we need to push back on this nonsense.
A funny story (well, funny because it didn't happen to me) was chronicled on a recent episode of Corey Quinn's "Whiteboard Confessions" -- https://www.lastweekinaws.com/podcast/aws-morning-brief/whit... -- where Slackbot's auto URL unfurling feature "clicked" an SNS alert unsubscribe link when a tech posted the a report including that link into a slack channel. If nobody lost their job over this triggering a SEV-1 al…
In practice it's going to be difficult because:
1) You can't do effective geolocation on cell phone IP address, except perhaps at a country level [1]
2) Your ex would need to open your message to trigger the link preview loading. If you're stalking them, they're probably either ignoring you or blocking you
3) At most, if the ex is connected to a Wi-Fi network like Starbucks and opens the message, you'll be able to get the city they're in, maybe.
But at the end of the day, what you'd get from link previews is no different from embedding a tracker image in an e-mail. And while it requires someone to suspect they're being stalked, blocking would prevent all of this.
[1] http://www.cs.yale.edu/homes/mahesh/papers/ephemera-imc09.pd...
Re: A Facebook crawler was making 7M requests per day to my stupid website
#67I just thought of a malicious idea. If I were Amazon or some other cloud provider, I could hammer my customers' web sites with tons of requests to make them pay more for resource usage (network bandwidth, s3 calls, misc per/unit usage etc). It would be hard to trace as well. Wonder if people are already doing that today.
Sometimes that doesn't stop the pointy haired bosses, because sometimes they are the perfect storm of unethical and moronic.
Re: A Facebook crawler was making 7M requests per day to my stupid website
#68Doesn't ignoring headers and taking someones site down fall afoul of the CFAA? Especially given how comments here are showing that this is a recurring issue? Depending on which side of the fence you're on, there could be standing to go after FB for this either for money, or for their lawyers to help set a precedent to limit CFAA further.
Whether or not the above even holds today (hiQ Labs v. LinkedIn argues it does not) remains to be decided in a current appeal to the Supreme Court.
Re: A Facebook crawler was making 7M requests per day to my stupid website
#69Re: A Facebook crawler was making 7M requests per day to my stupid website
#70Hate to say but this is probably a FB engineer running a test/experiment and not a production crawler taking robots.txt etc into account.