Live data from Hacker News

Facebook was used as a proxy by web scraping bots

datadome.co

91–100 of 125 posts

Re: Facebook was used as a proxy by web scraping bots

#91
Would just like to give an honorable mention to Google Translate, the most accessible http proxy of all time. It’s especially good for bypassing corporate access controls. I’ve used it many times for accessing solution threads on technical subreddits at work.

Re: Facebook was used as a proxy by web scraping bots

#92

Earlier quoted context omitted.

I'd place that blame towards website owners. Both Facebook and Twitter are pretty open where they read that info from, and an owner can pretty easily pass those fields (it's just some tags in the element). They also have their own validators: https://cards-dev.twitter.com/validator and https://developers.facebook.com/tools/debug/ The only issue I'm aware of is that Facebook's crawler breaks about every two months or…

What meta tags do I have to fill and why is Twitters/FBs preview suddenly my problem? > https://developer.twitter.com/en/docs/twitter-for-websites/c... So, I should have to include twitter specific meta tags even though I personally don't care about twitter? Maybe twitter should make it clear which tags they read? Maybe it's SEO bullshit I don't care about? Maybe even even the OG: tags don't work all the time and res…

If you don't want to fill them out, don't... Filling them out lets you customize your link preview on twitter. If you don't care about Twitter, why would this affect you at all?

Re: Facebook was used as a proxy by web scraping bots

#93
post #37

You can't create your own link previewer, cloudflare will put a captcha in front of every website. All I want is a a freaking tag. They don't seem eager to fix it either, their proposed solution is to contact every website owner (seriously) to ask them to whitelist you[1]. Frankly, i wish facebook or cloudflare offered their previewer as a free service, since most websites have them whitelisted. 1. https://community.…

I’ve long said Cloudflare is a dangerous threat to the open internet and as well as some privacy tools like TOR. But it doesn’t always get much traction on here because both the founder and employees of cloudflare are quite popular users on HN. Some have given me brief half assed counter answers that conveniently miss other harder questions like a good PR person does (and which you seem to have gotten in your reply).…

Hey, try https://deflect.ca if you want an ethical DDoS protection service.

Re: Facebook was used as a proxy by web scraping bots

#94

Earlier quoted context omitted.

> DRM on Video has not really caught on, at least on-line This seems like a weird statement. All of the paid streaming services use DRM on Video, so all major browsers include the requisite black-box DRM modules. I'm actually surprised YouTube has not added Widevine DRM for all videos yet, but I'm sure it'll happen if RIAA/MPAA get annoyed enough with youtube-dl and the like. > PS - By AMP, do you mean Amazon Prime?…

I've never heard of this AMP thing. Is it really that popular? Valid point about paid streaming services - which I don't use.

AMP is incredibly popular. Every news site has enabled it. You have a 100% chance of seeing an AMP page in the top results for anything.

Google had 2 options - make websites faster the normal way (remove bloat) or make websites faster by introducing AMP. AMP is controlled by Google. What do you think they did? They said they would reduce the site's ranking if they didnt use AMP. Within weeks, everybody except Wikipedia was introducing AMP.

Re: Facebook was used as a proxy by web scraping bots

#95

Earlier quoted context omitted.

To add to sibling - running your own mail server is the only way to ensure your email is not read by someone else which is so messed up.

> running your own mail server is the only way to ensure your email is not read by someone else But any mail you send to someone else probably ends up read by Google/Microsoft anyway, since that's where their mailbox is. Also, email security is a joke. It's 2020, and even TLS encrypted SMTP connections tend not to check for a valid certificate, making them trivial to MITM.

Practically speaking how does one MITM an SMTP connection? For example, from Google to Microsoft. They connect directly to the IP addresses they get from MX records + lookup. What's the actual threat vector/execution here?

Re: Facebook was used as a proxy by web scraping bots

#96

Earlier quoted context omitted.

I'm guessing that in these countries the cost of Internet traffic is dominated by the undersea cables at their borders. Facebook has been paying to have new undersea cables laid. This is done as part of a consortium, but those cables only have 6-12 strands in them (the repeaters are bulky) so owning even just one whole strand of fiber in an undersea cable is still an obscene amount of bandwidth for a single company t…

In The Philippines, my understanding is that they have ample bandwidth via Korea and other countries in the region. But the reason they have such expensive terrible internet is because of a lack of net neutrality and deregulation. The cellphone duopoly sells "YouTube passes", that entitle you to get unthrottled YouTube for brief periods of time.

Net neutrality isn’t related to Internet speeds. Good speeds are just driven by having competition.

Comcast was suddenly able to provide 1gbps for the same price as an 80mbps package when a fiber competitor entered the market.

Even with net neutrality, there is no incentive to make the internet better as an operator if you’re operating in a government granted monopoly/duopoly market.

Re: Facebook was used as a proxy by web scraping bots

#97
post #37

You can't create your own link previewer, cloudflare will put a captcha in front of every website. All I want is a a freaking tag. They don't seem eager to fix it either, their proposed solution is to contact every website owner (seriously) to ask them to whitelist you[1]. Frankly, i wish facebook or cloudflare offered their previewer as a free service, since most websites have them whitelisted. 1. https://community.…

I’ve long said Cloudflare is a dangerous threat to the open internet and as well as some privacy tools like TOR. But it doesn’t always get much traction on here because both the founder and employees of cloudflare are quite popular users on HN. Some have given me brief half assed counter answers that conveniently miss other harder questions like a good PR person does (and which you seem to have gotten in your reply).…

Normally skip those sites that ask for a Cloudflare captcha if the site isn't too important. Luckily this is the case most of the time.

Would be annoying when online banking or governmental sites start asking for them.

Re: Facebook was used as a proxy by web scraping bots

#98

Earlier quoted context omitted.

In The Philippines, my understanding is that they have ample bandwidth via Korea and other countries in the region. But the reason they have such expensive terrible internet is because of a lack of net neutrality and deregulation. The cellphone duopoly sells "YouTube passes", that entitle you to get unthrottled YouTube for brief periods of time.

Net neutrality isn’t related to Internet speeds. Good speeds are just driven by having competition. Comcast was suddenly able to provide 1gbps for the same price as an 80mbps package when a fiber competitor entered the market. Even with net neutrality, there is no incentive to make the internet better as an operator if you’re operating in a government granted monopoly/duopoly market.

Net neutrality eliminates the ability of an operator to discriminate and offer uncapped data or higher speed passes to just YouTube.

Re: Facebook was used as a proxy by web scraping bots

#99
post #37

Earlier quoted context omitted.

I’ve long said Cloudflare is a dangerous threat to the open internet and as well as some privacy tools like TOR. But it doesn’t always get much traction on here because both the founder and employees of cloudflare are quite popular users on HN. Some have given me brief half assed counter answers that conveniently miss other harder questions like a good PR person does (and which you seem to have gotten in your reply).…

> But it doesn’t always get much traction on here because both the founder and employees of cloudflare are quite popular users on HN. I don't think it gets much traction because you're barking up the wrong tree. Also, suggesting that YC is out to silence you and that nobody actually has a counter argument isn't very good for traction, either. Until my website can't get taken off by a $5 rental of an internet-of-shit…

I hosted a server that was attacked all the time over a comcast connection and was always able to figure it out without cloudflare proxy blocking for me

Re: Facebook was used as a proxy by web scraping bots

#100

You can't create your own link previewer, cloudflare will put a captcha in front of every website. All I want is a a freaking tag. They don't seem eager to fix it either, their proposed solution is to contact every website owner (seriously) to ask them to whitelist you[1]. Frankly, i wish facebook or cloudflare offered their previewer as a free service, since most websites have them whitelisted. 1. https://community.…

At https://host.io we scrape every registered domain once a month, and make the meta data available freely over an API. You could use that to get a title for a domain (although not for a URL that's not the main domain), eg:

    $ curl https://host.io/api/web/facebook.com?token=$TOKEN
    {
      "domain": "facebook.com",
      "rank": 2,
      "url": "https://www.facebook.com/",
      "ip": "157.240.11.35",
      "date": "2020-08-26T17:39:17.981Z",
      "length": 160817,
      "encoding": "utf8",
      "copyright": "Facebook © 2020",
      "title": "Facebook - Log In or Sign Up",
      "description": "Create an account or log into Facebook. Connect with friends, family and other people you know. Share photos and videos, send messages and get updates.",
      "links": [
        "messenger.com",
        "oculus.com"
      ]
   }
See https://host.io/docs for more details about the API and what else you can do with it (eg. find backlinks to domains, domains with the same adsense ID etc)
Post reply on HN