Live data from Hacker News

A Facebook crawler was making 7M requests per day to my stupid website

coding.napolux.com

221–230 of 416 posts

Re: A Facebook crawler was making 7M requests per day to my stupid website

#221
post #106

Earlier quoted context omitted.

Never forget Aaron Swartz. All he did was send download requests to JSTOR. https://www.youtube.com/watch?v=9vz06QO3UkQ

I'm going to get a ton of hate for this, but he was doing it with the intent to redistribute the content for free. He didn't own the content. There's a big difference. That being said it's very sad what came about of that. I really don't think the FBI needed to be involved.

Yeah and Facebook did this crawlings with intent to commit CFAA - computer fraud and abuse, now someone call the feds quickly before the perpetrators are zucked into hiding

Re: A Facebook crawler was making 7M requests per day to my stupid website

#222

Earlier quoted context omitted.

Next time, have some courtesy for the author and the rest of us by requesting via personal exchange over email instead of hijacking the thread and distracting from the conversation.

Boo hiss. There’s plenty of room here for a polite exchange between professionals. Collapse the thread and move on.

Come on. I love Ruud's posts and I took inspiration for my blog. He was right asking for a link, which I gladly provided.

That's it for me.

Re: A Facebook crawler was making 7M requests per day to my stupid website

#223
post #216
post #29

We've had the same issue. They were doing huge bursts of tens of thousands of requests in very short time several times a day. The bots didn't identify as FB (used "spoofed" UAs) but were all coming from FB owned netblocks. I've contacted FB about it, but they couldn't figure out why this was happening and didn't solve the problem. I found out that there is an option in the FB Catalog manager that lets FB auto-remove…

I dont really understand what is the issue. On my welcome page (while all other urls are impossible to guess) i give browser something that requires a few seconds of cpu at 100% to crunch. And tracking some user action in between, visting tarpitted urls etc. In last few years no bot came through. Why bother with robots.txt, just give them something to break their teeths... (I would give you the url, but I just dont w…

Sounds like terrible UX! Won't browsers give you a "Stop script" prompt if you do that?

Re: A Facebook crawler was making 7M requests per day to my stupid website

#224

Earlier quoted context omitted.

This pattern needs to die. Whenever I paste a link in any sort of messaging app - iMessages to Slack - it puts a thumbnail and summary, polluting the entire conversation with tons of noise. Messaging apps have no business in looking up the URL. Just let it pass as a link. Fuck everything about this and we need to push back on this nonsense.

I like link previews. I don't like when they are managed server-side instead of client-side.

Also considering the sibling story, at least the server-side preview builder is definitely entirely unauthenticated to every possible service. A preview builder running on the client side might conceivably pick up your authentication to something in various circumstances.

Re: A Facebook crawler was making 7M requests per day to my stupid website

#225
post #202

Earlier quoted context omitted.

Next time, have some courtesy for the author and the rest of us by requesting via personal exchange over email instead of hijacking the thread and distracting from the conversation.

> ... and the rest of us ... Please don't try to police the thread and speak for yourself only.

“And the rest of us, excluding Dahoon”

Re: A Facebook crawler was making 7M requests per day to my stupid website

#226
post #189

Are you sure it's a crawler and not a proxy of some sort? Eg. one of your links is on something high traffic in facebook, and all requests are human, running through fb machines.

Nah, it's a specific user-agent for their crawler

Re: A Facebook crawler was making 7M requests per day to my stupid website

#227

Earlier quoted context omitted.

So, just to be clear, you are perfectly fine with someone taking something someone else created and violating the terms upon which they were given that thing? e.g. If I took code you wrote and lets say released under an MIT license and claimed I wrote it and didn't give you any credit, and in fact released it under another license entirely, you'd be fine with that?

It's perfectly valid to criticize the original license choice. GPLv3 is a very restrictive license, especially for what is essentially a micro blog (though I dislike the license for most open source software anyway). Add on the original author going after a bit of CSS, not even the main effort of the project in question, and you've got my "petty" comment.

There is an artist who takes images from magazines, repurposes them for his own art, and sells them for hundreds of thousands of dollars, then gets sued by the magazines & photographers & artists, and guess what, HE WINS AGAINST THOSE LAWSUITS. Copyright is BS, especially concerning HTML & CSS code. What a joke.

Re: A Facebook crawler was making 7M requests per day to my stupid website

#228
post #129

Earlier quoted context omitted.

Super optimized in a SEO sense, not in a pleasing the HN crowd kind of sense I'd assume ;)

You're right, but for some reason the whole SEO thing just winds me up. It's my opinion that 'good' SEO makes sites worse for actual people to use.

SEO is for machines, not for users IMHO.

Re: A Facebook crawler was making 7M requests per day to my stupid website

#229
post #140

Hi Napolux, It looks like your site is using a theme based on my website ( https://ruudvanasseldonk.com/ , source at https://github.com/ruuda/blog ). That is fine — it is open source after all, licensed under the GPLv3. But I can’t find the source code for your site, and I can’t find any prominent notices saying that you modified my source. Could you please add those?

Your request is extremely petty and you should be thoroughly ashamed of yourself, give up on your career as a web developer, and do something else with your life.

Re: A Facebook crawler was making 7M requests per day to my stupid website

#230
post #167

Earlier quoted context omitted.

Sure man no problem. The code for my theme is here BTW with credits To your original blog https://github.com/napolux/coding.napolux.com/

Thanks, I’m flattered to see it be used as inspiration :) I searched quickly but I didn’t find that repository. You might want to link it somewhere in your footer or from a comment in the html.

Flattered that your few, boring lines of CSS and HTML (that could easily be reproduced by a monkey) was used on someone else's website?
Post reply on HN