isn't this what robots.txt for? can't you block facebook?
As the article clearly states: no
A Facebook crawler was making 7M requests per day to my stupid website
101–110 of 416 posts
Re: A Facebook crawler was making 7M requests per day to my stupid website
#102How do you monetize a robot?
That's a good question. Even better if Facebook can pay me some money :P
Re: A Facebook crawler was making 7M requests per day to my stupid website
#103isn't this what robots.txt for? can't you block facebook?
Re: A Facebook crawler was making 7M requests per day to my stupid website
#104Earlier quoted context omitted.
Hey! Facebook engineer here. If you have it, can you send me the User-Agent for these requests? That would definitely help speed up narrowing down what's happening here. If you can provide me the hostname being requested in the Host header, that would be great too. I just sent you an e-mail, you can also reply to that instead if you prefer not to share those details here. :-)
I'm not sure I'd publicly post my email like that, if I worked at FB. But congratulations on your promotion to "official technical contact for all facebook issues forever".
Re: A Facebook crawler was making 7M requests per day to my stupid website
#105Earlier quoted context omitted.
When someone does this to Facebook its malicious and they go to jail. WHen Facebook does it to someone else... "oops".
From the point of view of the law, the intent (malicious or benevolent) is often as important as the action itself.
Re: A Facebook crawler was making 7M requests per day to my stupid website
#106Earlier quoted context omitted.
When someone does this to Facebook its malicious and they go to jail. WHen Facebook does it to someone else... "oops".
Never forget Aaron Swartz. All he did was send download requests to JSTOR. https://www.youtube.com/watch?v=9vz06QO3UkQ
Re: A Facebook crawler was making 7M requests per day to my stupid website
#107Earlier quoted context omitted.
Never forget Aaron Swartz. All he did was send download requests to JSTOR. https://www.youtube.com/watch?v=9vz06QO3UkQ
I'm going to get a ton of hate for this, but he was doing it with the intent to redistribute the content for free. He didn't own the content. There's a big difference. That being said it's very sad what came about of that. I really don't think the FBI needed to be involved.
There is actually no proof of this.
Re: A Facebook crawler was making 7M requests per day to my stupid website
#108Shouldn't anybody doing such thing be liable, and be sued for negligence and required to pay damages?
That sounds like a lot of bandwidth (and server stress).
Re: A Facebook crawler was making 7M requests per day to my stupid website
#109Earlier quoted context omitted.
I'm going to get a ton of hate for this, but he was doing it with the intent to redistribute the content for free. He didn't own the content. There's a big difference. That being said it's very sad what came about of that. I really don't think the FBI needed to be involved.
> he was doing it with the intent to redistribute the content for free There is actually no proof of this.
I have no idea where I got that impression from, but I do recall reading several articles about him, as well as his blog around that time. It's possible I just internalized others' assumptions about his behavior, but it certainly seemed like an in-character thing for him to to.
Re: A Facebook crawler was making 7M requests per day to my stupid website
#110Make sure you file a bug - there are a myriad of sources internally, but we can often hunt it down easily enough (assuming it gets triaged to eng). Important info is the host of the urls being crawled, and the User Agent attached to the requests (headers are also good too). A timeseries graph of the hits (with date & timezone specified) can also help.
Why is this the owner's problem? Someone at Facebook should be filing the bug, or better yet instrumenting their systems so that incidents like this issue a wake-up-an-engineer alert. Fuck this culture of "it's up to the victim of our fuckup to file a bug report with us".