Live data from Hacker News

A Facebook crawler was making 7M requests per day to my stupid website

coding.napolux.com

171–180 of 416 posts

Re: A Facebook crawler was making 7M requests per day to my stupid website

#171

Earlier quoted context omitted.

Ask me anything. That specific website is a personal project of mine which I'd like not to disclose. As you can see from my blog I'm not selling anything, I don't even have a banner on my blog, so where your suspect is coming from?

Sorry, I was mostly joking. I definitely do believe Facebook could be responsible for this.

No problem :) I’m not the only one suffering this as you can see from the thread

Re: A Facebook crawler was making 7M requests per day to my stupid website

#172
post #163

Earlier quoted context omitted.

They look similar at a glance (border-top + Calluna font), so he might have taken inspiration from yours, but doesn't seem to have used any of your assets - the styles are clearly different and based on the WP 'BlankSlate' theme. (my personal blog had a top border like that a decade ago, when styles on the body were a novelty :))

I checked the source and it does indeed mention https://wordpress.org/themes/blankslate/ , but judging from the screenshot there, that theme is just really a blank theme with no style at all, and it was used to include a different stylesheet. The real style is at https://coding.napolux.com/wp-content/themes/coding.napolux.... , which looks like normalize.css followed by a Wordpress adaptation of my stylesheet. The si…

Come on, this is super petty. And that includes your license choice and your enforcement for this.

Re: A Facebook crawler was making 7M requests per day to my stupid website

#173
post #140

Hi Napolux, It looks like your site is using a theme based on my website ( https://ruudvanasseldonk.com/ , source at https://github.com/ruuda/blog ). That is fine — it is open source after all, licensed under the GPLv3. But I can’t find the source code for your site, and I can’t find any prominent notices saying that you modified my source. Could you please add those?

Next time, have some courtesy for the author and the rest of us by requesting via personal exchange over email instead of hijacking the thread and distracting from the conversation.

If it were a financial thing I'd agree but if you aren't complying with the GPL I think it's good to set an example (but not shaming)

Re: A Facebook crawler was making 7M requests per day to my stupid website

#174

It's an interesting thing to think about, that someone can make a request to your website and you pay for it. This is not how the mail works, where someone needs to buy a stamp to spam you.

It’s because we as an industry decided to give in and submit to the cloud providers’ bullshit model of paying overpriced amounts for bandwidth while good old bare-metal providers still offer unmetered bandwidth for very reasonable prices.

Re: A Facebook crawler was making 7M requests per day to my stupid website

#175
post #111

Did anyone notice the branded IP address? 2a03:2880:20ff:d::face:b00c "face:b00c"

FB has 2a00::/12 (and probably other blocks?) and can make something that looks like its company name from [0-9a-f]. Why wouldn't it do something harmless and fun like this? It's not as if it requires special dispensation or is breaking any rules.

I don't think the person was implying it's bad. Just pointing out an easter egg.

Re: A Facebook crawler was making 7M requests per day to my stupid website

#176
post #89
post #65

Earlier quoted context omitted.

Thanks man! I'll have a look.

Hey! Facebook engineer here. If you have it, can you send me the User-Agent for these requests? That would definitely help speed up narrowing down what's happening here. If you can provide me the hostname being requested in the Host header, that would be great too. I just sent you an e-mail, you can also reply to that instead if you prefer not to share those details here. :-)

Thanks for looking at this!

Re: A Facebook crawler was making 7M requests per day to my stupid website

#177

Earlier quoted context omitted.

Despite how it sounds, I ask this with zero judgment and pure curiosity. Why do you care?

The author probably cares about either copy-left or the right to tinker. Some people believe it's important and others don't. Basically any derivative product out of GPL3 code needs to either be open-source/copy-left, or if it uses the code as a library needs to let end users substitute that library for their own version.

Worse than unimportant, I find copyleft harmful to the developer community as a whole. Software is hardly "free" when you have to bend to the author's demands to license your entire rest of your project under their ideology if you want to use it, and I think we should stop calling it so. Maybe "Conditionally Free Software" instead. You're hardly helping the world releasing libraries in terms that nobody but hobbyists (who are going to rip your code out and replace it if they start making a proprietary product, which is not evil) will find acceptable. It works for the Linux kernel and self-contained applications but that's it.

Re: A Facebook crawler was making 7M requests per day to my stupid website

#178

Earlier quoted context omitted.

To me it pollutes the discussion away from the messages and puts a bunch of thumbnails where it could just be a clean text conversation. The thing that bothers me is also that I, as a user, have no choice - it just does it automatically. Do you like IRC?

Whatsapp allows to delete the link box: enter the link, wait for the preview, backspace, and send. The hyperlink remains clickable too.

Slack does it as well

Re: A Facebook crawler was making 7M requests per day to my stupid website

#179

It's an interesting thing to think about, that someone can make a request to your website and you pay for it. This is not how the mail works, where someone needs to buy a stamp to spam you.

It’s because we as an industry decided to give in and submit to the cloud providers’ bullshit model of paying overpriced amounts for bandwidth while good old bare-metal providers still offer unmetered bandwidth for very reasonable prices.

The question isn't about pricing. The fact that you pay for your own bandwidth and metal is a problem.

Re: A Facebook crawler was making 7M requests per day to my stupid website

#180
post #140

Hi Napolux, It looks like your site is using a theme based on my website ( https://ruudvanasseldonk.com/ , source at https://github.com/ruuda/blog ). That is fine — it is open source after all, licensed under the GPLv3. But I can’t find the source code for your site, and I can’t find any prominent notices saying that you modified my source. Could you please add those?

Next time, have some courtesy for the author and the rest of us by requesting via personal exchange over email instead of hijacking the thread and distracting from the conversation.

Boo hiss.

There’s plenty of room here for a polite exchange between professionals.

Collapse the thread and move on.

Post reply on HN