Live data from Hacker News

Taking action against scraping for hire

about.fb.com

221–230 of 240 posts

Re: Taking action against scraping for hire

#221
post #174

Earlier quoted context omitted.

The users agreed to share their data with Facebook, not some other company. If they didn't prevent this, they'd be asking for another Cambridge Analytica

That is a very good point, but surely it was taken into consideration when scraping was declared legal?

All that case says is "scraping is not a violation of the CFAA". But of course the scraped data still exists in legal limbo; maybe you can compute derived information from it, but the moment a scraper reproduces it there is all of copyright law waiting for them.

Re: Taking action against scraping for hire

#222
post #115

Earlier quoted context omitted.

Anything behind a login gate is private data for that registered user only.

Depends if another user can also access it, or whether the original author/owner of the data in question intends for it to be public. In Facebook's case, there are permission levels you can set on posts, including a "public" option (which isn't actually public though and will require a login anyway, but it can be any login) which would settle that debate quickly - hell I wouldn't be surprised if that option were to b…

> In Facebook's case, there are permission levels you can set on posts, including a "public" option (which isn't actually public though and will require a login anyway, but it can be any login)

Q: Have you tried this?

In a private browser session I started at google.com, searched for "site:facebook.com nextgrid", picked some random post, click through, and was reading the post without anything other than seeing FB's cookie banner. No sign of any login (which is good 'cause I don't have one)

Re: Taking action against scraping for hire

#223
post #38

Earlier quoted context omitted.

Yes. I want a free and open web.

Good for you. Normal people do not want posts shared privately amongst friends to become publicly available.

Then you need to trust your friends, because copy/paste and screenshots exist.

Re: Taking action against scraping for hire

#224

Collecting the rhetorical BS: "scraping attacks" Scraping is not an attack. Monopolists want to pretend they own your data because they get unlimited access to monetize it whereas competitors should have none. "self-compromised" Monopolists want to sell you thus it's imperative they maintain the fiction of "one person, one account". By admitting you own your account, they'd have to allow sharing and they wouldn't be…

[dead]

Re: Taking action against scraping for hire

#225

Collecting the rhetorical BS: "scraping attacks" Scraping is not an attack. Monopolists want to pretend they own your data because they get unlimited access to monetize it whereas competitors should have none. "self-compromised" Monopolists want to sell you thus it's imperative they maintain the fiction of "one person, one account". By admitting you own your account, they'd have to allow sharing and they wouldn't be…

The only argument I have here (sadly in favor of FB) is with "safeguard people against clone sites". While I did give my data to FB, I didn't approve that transfer to another site/system. That is the only place I could possibly see some legal foot hold.

What happens when FB builds a shadow instagram profile of you based on your FB account? That already happens. FB clones their own data for other projects no different than what you might fear happening if this data were cloned to a third party. The cat is out of the bag already but FB wants to pretend they are the only ones with the right to abuse.

Re: Taking action against scraping for hire

#226

Collecting the rhetorical BS: "scraping attacks" Scraping is not an attack. Monopolists want to pretend they own your data because they get unlimited access to monetize it whereas competitors should have none. "self-compromised" Monopolists want to sell you thus it's imperative they maintain the fiction of "one person, one account". By admitting you own your account, they'd have to allow sharing and they wouldn't be…

Missed this one: > a US subsidiary of a "Chinese national" "high-tech" enterprise Replacing it with "a business" would do just fine.

No post body was provided.

Re: Taking action against scraping for hire

#227

Earlier quoted context omitted.

Depends if another user can also access it, or whether the original author/owner of the data in question intends for it to be public. In Facebook's case, there are permission levels you can set on posts, including a "public" option (which isn't actually public though and will require a login anyway, but it can be any login) which would settle that debate quickly - hell I wouldn't be surprised if that option were to b…

> In Facebook's case, there are permission levels you can set on posts, including a "public" option (which isn't actually public though and will require a login anyway, but it can be any login) Q: Have you tried this? In a private browser session I started at google.com, searched for "site:facebook.com nextgrid", picked some random post, click through, and was reading the post without anything other than seeing FB's…

I suspect it depends on your region, page/post in question and browser fingerprint. A post marked as public isn't 100% guaranteed to be publicly viewable. Sometimes you can view it but merely scrolling down on the page would trigger a login form for example (I've had this happen for pages that are definitely meant to be public such as businesses who'd have an interest in getting as many eyeballs as possible on their content).

I might be wrong and maybe the behavior is actually fully deterministic and isn't nefarious, but knowing the company behind it I'll assume malice until proven otherwise.

Re: Taking action against scraping for hire

#228

Earlier quoted context omitted.

Nah, you are straight-up wrong. In fact, it’s the opposite - the only companies who are scared of scraping are the ones whose business models rely on artificial lock-in, and we should all be working as hard as we can to demolish them.

>the only companies who are scared of scraping are the ones whose business models... This is just patently false. There is an expense incured by scraping. There is no benefit to a host providing the data from those scrapers. My logs are full of various bots that pull data from my webhost that costs me money to serve. I run various sites that do not serve ads. I do not include any 3rd party tracking. They're just simp…

Hey, I totally accept people have views other than my own. I just disagree with them.

It seems extremely weird that you’d want to publish content, but then get mad that people are using the thing that you published. But you do you.

Re: Taking action against scraping for hire

#229

Earlier quoted context omitted.

>the only companies who are scared of scraping are the ones whose business models... This is just patently false. There is an expense incured by scraping. There is no benefit to a host providing the data from those scrapers. My logs are full of various bots that pull data from my webhost that costs me money to serve. I run various sites that do not serve ads. I do not include any 3rd party tracking. They're just simp…

Hey, I totally accept people have views other than my own. I just disagree with them. It seems extremely weird that you’d want to publish content, but then get mad that people are using the thing that you published. But you do you.

How is that weird? I publish on my site to have people visit my site. I don't publsh for people to take my data and do what they will without attribution for where they got the data. How that makes no sense to others has me saying please don't do you because you are being not considerate to others

Re: Taking action against scraping for hire

#230

Earlier quoted context omitted.

The user agreed in facebook to have is data "public", so it can't complain that a robot scrap it. Nothing prevents him to restrict access to his pages an data to "trusted" friends.

The description in the article sounds like it scrapes private profile data. > Octopus designed the software to scrape data accessible to the user when logged into their accounts

I don't think so, it is more like you scrape what is accessible to this user. So in the end you will scrape your friends data. This is why I said that you are free to only share with friends that 'you trust'.
Post reply on HN