Live data from Hacker News

Taking action against scraping for hire

about.fb.com

211–220 of 240 posts

Re: Taking action against scraping for hire

#211
post #115

Earlier quoted context omitted.

Anything behind a login gate is private data for that registered user only.

> Anything behind a login gate is private data for that registered user only That's quite the claim, if only the login gate were either always there or indeed always not. Presuambly such "private" data ought not to be being indexed by search engines and returned to users who search? "site:instagram.com" is of the order of 228 million pages on google.com, and "site:facebook.com" is another 422 million.

pretty sure you get hit with a login gate if you navigate to the results via site:instagram.com no?

Re: Taking action against scraping for hire

#212

Earlier quoted context omitted.

Exactly. Your list of friends does not belong to Facebook, it belongs to you. I am sure Facebook believes they deserve a monopoly for having obtained it first. They do not. The market forces you to compete for every dollar you earn, so you have every right to expect Facebook to compete for every dollar they earn, and "I touched it first therefore it's mine!" is not competition.

But, but, but..... you agreed that Facebook does own your friends list when you signed up for an account and started giving them all your data. If I run a restaurant, and I stipulate that when you walk through the doors and place an order I reserve the right to take your picture and post it on the bulletin board, why would you place the order and then get pissed off when I post a picture of you on the bulletin board?…

If your bulletin board somehow let you monopolize the restaurant industry (? lol) then we should absolutely vote for some politicians to boot your entitled ass back into competition.

Obviously, the idea of a bulletin board granting a restaurant an effective monopoly is ridiculous so your analogy is trash, but even if your analogy wasn't trash, your conclusion would still be wrong.

Re: Taking action against scraping for hire

#213

From GDPR point-of-view this kind of 3rd party data collection is not acceptable (assuming it covers personal information, for example names of people and what they have posted). The difference with Meta's own data collection is that the users have relationship with Meta and users have given their permission for Meta to handle the data. Users also know they can contact Meta and ask them to remove the data. 3rd partie…

From a GDPR point of view the scraper would be acting as a data processor on behalf of their customer, no different from using a cloud storage service for your contacts. It's fine as long as the third-party doesn't misuse the scraped data or share it with third-parties and there's no evidence they did so in this case.

> and there's no evidence they did so in this case.

Indeed; the users probably wanted to make the data public, if scraper accounts could see it. There is a GDPR allowance for data "manifestly made public by the data subject".

https://gdpr-info.eu/art-9-gdpr/

Here, it's just Facebook wanting to keep the data inside a walled garden.

For the same reason, I quit LinkedIn and made my own site. I don't want people to have to sign in to see my profile.

Re: Taking action against scraping for hire

#214

Collecting the rhetorical BS: "scraping attacks" Scraping is not an attack. Monopolists want to pretend they own your data because they get unlimited access to monetize it whereas competitors should have none. "self-compromised" Monopolists want to sell you thus it's imperative they maintain the fiction of "one person, one account". By admitting you own your account, they'd have to allow sharing and they wouldn't be…

Missed this one:

> a US subsidiary of a "Chinese national" "high-tech" enterprise

Replacing it with "a business" would do just fine.

Re: Taking action against scraping for hire

#215
post #210

Earlier quoted context omitted.

Former attorney turned software developer here! Nope, it's not a settled question in the way that I think you mean. Each ToS is different so each would be subject to individual legal analysis in court on its own terms. Questions would include whether the ToS is unconscionable, whether the terms violate laws of the locality/nation, and so forth. It's the same with traditional contracts - the fact that contracts have b…

Why can't FB simply include a clause like "No kind of automated scraping is allowed, except for search engines in robots.txt"? This would save them so much time in court, arguing over the use of fake accounts which should really be irrelevant.

It's not clear that clause would be enforceable. Scraping has been found to be lawful in many jurisdictions, including the US, even without the consent of the host.

Re: Taking action against scraping for hire

#217
post #57

Earlier quoted context omitted.

This is a false expectation and it’s important people learn this.

They’ll stop posting in the way they currently enjoy and will, therefore, have lost some freedom. Great outcome! In other news: your partner may also leak your most intimate secrets. I hope they do, to teach you a lesson? Every trust can be betrayed. Why do you believe a world without trust would be better? Only because you cannot handle the nuance of different levels of trust?

The freedom to live in a fictional world where Facebook safeguards your data is just as available regardless the reality of the situation.

The reality of the situation is that Facebook is a walled garden built on the labor of it's users and it is objecting to those users reclaiming the fruits of their labor by scraping.

Re: Taking action against scraping for hire

#219

“This industry makes scraping available to individuals and companies that otherwise would not have the capabilities.” - seems like web scraping companies are doing a good job :)

Maybe some irony here as IIRC Facebook started as essentially a scraping company, pulling student profiles from college websites and re-publishing it for their own profit.

The scrapers have become the scrapees. The horror.

Re: Taking action against scraping for hire

#220
post #211

Earlier quoted context omitted.

> Anything behind a login gate is private data for that registered user only That's quite the claim, if only the login gate were either always there or indeed always not. Presuambly such "private" data ought not to be being indexed by search engines and returned to users who search? "site:instagram.com" is of the order of 228 million pages on google.com, and "site:facebook.com" is another 422 million.

pretty sure you get hit with a login gate if you navigate to the results via site:instagram.com no?

> you get hit with a login gate if you navigate to the results via site:instagram.com

Nope, I just tried it (private browser session, no IG activity from my IP recently)

google.com -> "site:instagram.com nojito" -> results -> www.instagram.com/explore/tags/nojito/ with a page of photos.

Quickly scrolling down the page for several dozen photos does eventually trigger the login box, though.

Post reply on HN