Live data from Hacker News

Facebook was used as a proxy by web scraping bots

datadome.co

61–70 of 125 posts

Re: Facebook was used as a proxy by web scraping bots

#61
post #33

At the best case scenario, google has a monopoly on scraping. Imagine trying to create a global search engine, how can you possibly even crawl sites that are behind cloudflare or just allow google/fb/bing bots ? Can you real-time crawl twitter ? Pretty sure they have a special deal with google to instant ping on new tweets. How many websites actually ping google on new content ?

This is so fucked. We've encrusted ourselves into this walled fiefdom, and there's no way to break free. It's only going to get worse. Chrome will be the only browser. AMP the only delivery mechanism. Video will require DRM. Eventually, text content will too. Binary blobs with no ad blocking.

Well, I think you're being overly alarmist, at least in the short term: DRM on Video has not really caught on, at least on-line; Non-Chrome browsers continue to have a significant share (mostly on Desktops); and ad blocking remains rather effective.

In the long run, I'm definitely worried: Capitalist economies tend to see a concentration of capital, generally and in most sectors individually. And this seems to be a real danger with computing technology. Coupled with mass surveillance and the pushing of people to have their personal information held by those large tech companies, a dystopia is not inconceivable.

PS - By AMP, do you mean Amazon Prime?

Re: Facebook was used as a proxy by web scraping bots

#62

Earlier quoted context omitted.

I don't get what value link previews add. Someone shares a link with me (on skype, slack, teams... whatever) and I care about the content because the person sharing it with me thinks I could/should care about it, or someone shares a link on an aggregator and then I don't think it is too much to ask for that someone to write a summary. If the link is worth sharing writing 1 sentence to explain why isn't too much to as…

They're sending you traffic. Imagine Twitter or Facebook without link preview, it's much harder to use and overall reduces the change I'll click on a link. Do you think only Twitter and Facebook should be allowed publish previews?

>They're sending you traffic.

Irrelevant traffic for every metric I care about.

>Imagine Twitter or Facebook without link preview

That's exactly what I'm saying. Either I care about what that person thinks might interest me or I don't. The link preview abstract is shit anyway. Does the site title and the 2 sentence abstract really sway you? If someone wants to send traffic my way, writing an interesting abstract is not too much to ask.

>it's much harder to use and overall reduces the change I'll click on a link

Maybe you should re-evaluate who you follow on twitter? I frankly could care less about facebook.

>Do you think only Twitter and Facebook should be allowed publish previews?

I think previews are worthless regardless, I thought I made that clear. Either you care about me linking it to you or you do not.

*EDIT: And just for fun, here is the link preview stuff from my latest skype call with my brother: https://imgur.com/a/yO5OP36

Look at all the value those previews added.

Re: Facebook was used as a proxy by web scraping bots

#63

You can't create your own link previewer, cloudflare will put a captcha in front of every website. All I want is a a freaking tag. They don't seem eager to fix it either, their proposed solution is to contact every website owner (seriously) to ask them to whitelist you[1]. Frankly, i wish facebook or cloudflare offered their previewer as a free service, since most websites have them whitelisted. 1. https://community.…

I don't get what value link previews add. Someone shares a link with me (on skype, slack, teams... whatever) and I care about the content because the person sharing it with me thinks I could/should care about it, or someone shares a link on an aggregator and then I don't think it is too much to ask for that someone to write a summary. If the link is worth sharing writing 1 sentence to explain why isn't too much to as…

when you paste a link on reddit and it autocompletes the title

update a bookmark title, or check if it exists.

is it not self-evident that a link being crawlable is useful?

Re: Facebook was used as a proxy by web scraping bots

#64
post #37

Earlier quoted context omitted.

I’ve long said Cloudflare is a dangerous threat to the open internet and as well as some privacy tools like TOR. But it doesn’t always get much traction on here because both the founder and employees of cloudflare are quite popular users on HN. Some have given me brief half assed counter answers that conveniently miss other harder questions like a good PR person does (and which you seem to have gotten in your reply).…

Edit: my bad. Misinterpreted your comment. Can you elaborate on how Tor is a threat to the open internet? That's a non-obvious statement to me. I'm aware that it's compromisable via controlling exit nodes (NSA, various nations) but that's not really the threat profile for the average person. Are there any other reasons? Because despite its flaws, afaik TOR is an attempt to make the internet _more_ open to those who a…

Website owners can actually whitelist Tor traffic as a "country", but not a lot of them knows/cares/wants to do that.

Re: Facebook was used as a proxy by web scraping bots

#65

Earlier quoted context omitted.

I don't get what value link previews add. Someone shares a link with me (on skype, slack, teams... whatever) and I care about the content because the person sharing it with me thinks I could/should care about it, or someone shares a link on an aggregator and then I don't think it is too much to ask for that someone to write a summary. If the link is worth sharing writing 1 sentence to explain why isn't too much to as…

when you paste a link on reddit and it autocompletes the title update a bookmark title, or check if it exists. is it not self-evident that a link being crawlable is useful?

>when you paste a link on reddit and it autocompletes the title

Oh no, you have to copy/paste the title?

>update a bookmark title, or check if it exists.

I can access the site without a captcha, my browser can fetch the title.

>is it not self-evident that a link being crawlable is useful?

No, it is not. Maybe a site owner does not want crawlers to index the site?

Me being able to access the title and any html meta tags is not the same as some crawler being able to access it. It seems like your beef is with cloudflare and that is fine but please state that that is your issue and don't try to frame it as something else. What I don't get is how everybody places the blame at cloudflares feet. It is my choice as a host to use cloudflare and to use their protection features.

Re: Facebook was used as a proxy by web scraping bots

#66

Earlier quoted context omitted.

You're really pulling a "how hard could it really be??" to DDoS prevention? You should at least be humbled by how few services can even offer DDoS protection that works against volumetric attacks and isn't just based on null-routing. The people with skin and money in the game might know something you don't.

here's how simple it is : if (!website.underDDoS && website.requestedTimesToday[ip]

How do you implement "website.underDDoS"?

Through a proxy - mind you; CloudFlare makes their decision without access to your CPU or DB metrics, and don't know which page load times are legitimately slow and which aren't supposed to be.

Re: Facebook was used as a proxy by web scraping bots

#67

Earlier quoted context omitted.

They're sending you traffic. Imagine Twitter or Facebook without link preview, it's much harder to use and overall reduces the change I'll click on a link. Do you think only Twitter and Facebook should be allowed publish previews?

Half the time the link preview picks the wrong picture and sometimes even the quote. Twitter and Facebook would both be improved by disabling it. Hell, it might even stop people from thinking they need a hero image for their 2 paragraph medium shitpost.

I'd place that blame towards website owners. Both Facebook and Twitter are pretty open where they read that info from, and an owner can pretty easily pass those fields (it's just some tags in the element).

They also have their own validators: https://cards-dev.twitter.com/validator and https://developers.facebook.com/tools/debug/

The only issue I'm aware of is that Facebook's crawler breaks about every two months or so.

Re: Facebook was used as a proxy by web scraping bots

#68
post #33

Earlier quoted context omitted.

This is so fucked. We've encrusted ourselves into this walled fiefdom, and there's no way to break free. It's only going to get worse. Chrome will be the only browser. AMP the only delivery mechanism. Video will require DRM. Eventually, text content will too. Binary blobs with no ad blocking.

Well, I think you're being overly alarmist, at least in the short term: DRM on Video has not really caught on, at least on-line; Non-Chrome browsers continue to have a significant share (mostly on Desktops); and ad blocking remains rather effective. In the long run, I'm definitely worried: Capitalist economies tend to see a concentration of capital, generally and in most sectors individually. And this seems to be a r…

> DRM on Video has not really caught on, at least on-line

This seems like a weird statement. All of the paid streaming services use DRM on Video, so all major browsers include the requisite black-box DRM modules. I'm actually surprised YouTube has not added Widevine DRM for all videos yet, but I'm sure it'll happen if RIAA/MPAA get annoyed enough with youtube-dl and the like.

> PS - By AMP, do you mean Amazon Prime?

I think he means Google AMP[1], which is slowly infecting more of the top search results on Google.

[1] https://developers.google.com/amp/

Re: Facebook was used as a proxy by web scraping bots

#69

Earlier quoted context omitted.

Half the time the link preview picks the wrong picture and sometimes even the quote. Twitter and Facebook would both be improved by disabling it. Hell, it might even stop people from thinking they need a hero image for their 2 paragraph medium shitpost.

I'd place that blame towards website owners. Both Facebook and Twitter are pretty open where they read that info from, and an owner can pretty easily pass those fields (it's just some tags in the element). They also have their own validators: https://cards-dev.twitter.com/validator and https://developers.facebook.com/tools/debug/ The only issue I'm aware of is that Facebook's crawler breaks about every two months or…

What meta tags do I have to fill and why is Twitters/FBs preview suddenly my problem?

>https://developer.twitter.com/en/docs/twitter-for-websites/c...

So, I should have to include twitter specific meta tags even though I personally don't care about twitter? Maybe twitter should make it clear which tags they read? Maybe it's SEO bullshit I don't care about? Maybe even even the OG: tags don't work all the time and result in dumb previews?

Re: Facebook was used as a proxy by web scraping bots

#70
post #21

Earlier quoted context omitted.

Because if you don't have it some a-hole will go and ddos your site or you want to prevent a hug-of-death because of reasons. It seems a lot of issues happen because bad players are continued to allowed to thrive, example: everybody uses a big provider because they're the only ones that solved the spam issue.

I use Zoho.com and I rarely get spam, if ever.

I use it as well and I get sooo much more spam than I git on Gmail.
Post reply on HN