Live data from Hacker News

Taking action against scraping for hire

about.fb.com

61–70 of 240 posts

Re: Taking action against scraping for hire

#61
post #15

>Octopus, a US subsidiary of a Chinese national high-tech enterprise, built a cloud-based platform designed to provide paying customers access to on-demand scraping software and services. It is interesting as how they try to position this as a Chinese attack on them.

it look like Zack is giving up on the Chinese market.

Re: Taking action against scraping for hire

#62
post #38

Earlier quoted context omitted.

Yes. I want a free and open web.

Good for you. Normal people do not want posts shared privately amongst friends to become publicly available.

Then why would you ever put it on a website that generates its revenue from using and selling your data?

Re: Taking action against scraping for hire

#63
I'm torn on Web scraping because the extreme of each end of the spectrum on this issue both seem unreasonable.

On one side, you have people who say any form of scraping is be disallowed, even prosecutable. This went so far that the Department of Justice on behalf of AT&T prosecuted a case of URL modification [1]. One of the few bright spots for this psychotic Supreme Court was to curtail the government's power under the CFAA by limiting what constituted "unauthorized" access [2].

On the other hand, there are those who think that any level of scraping should be fine and I think that's untenable too. Consider Yahoo indexing of Stack Overflow [3]:

> In the meantime, since Yahoo (via Slurp!) is about 0.3% of our traffic, but insists on rudely consuming a huge chunk of our prime-time bandwidth, they’re getting IP banned and blocked.

Do these "scraping extremists" think such actions should be illegal? It's actually not that far-fetched given the Ninth Circuit decided LinkedIn wrongly blocked HiQ scraping [4]. Like if you change your website with the intent that it'll make scraping more difficult, is that a problem? What if it's an unintended side effect?

Additionally, companies like Meta, Google and Apple are going to be way more acountable to abiding by data retention laws and regulations than any scraper. If it's OK to scrape FB.com completely, that information is out there forever.

I certainly think the government shouldn't prosecute on behalf of companies. At least that should expose to people how the government's #1 priority is in fact to protect the true constituents: corporations and the capital-owning class.

[1]: https://www.techdirt.com/2013/09/30/dojs-insane-argument-aga...

[2]: https://en.wikipedia.org/wiki/Van_Buren_v._United_States

[3]: https://stackoverflow.blog/2009/06/16/the-perfect-web-spider...

[4]: https://blog.ericgoldman.org/archives/2019/09/ninth-circuit-...

Re: Taking action against scraping for hire

#65

Of course, Facebook wants to make it sound like scraping is illegal, when it generally isn't. But account hijacking and mass-creation of accounts just to access private pages are clear violations of the Facebook and Instagram ToS, so they surely can sue for that.

[deleted]

Re: Taking action against scraping for hire

#66
post #13

This is different from LinkedIn v HiQ because HiQ was only scraping publicly available data that was generally accessible to the broader internet. In these two cases, the data is being scraped from FB/Insta using credentials that the client handed over or the mass creation of accounts solely for scraping purposes.

Yeah, I think this is more like the Cambridge Analytica situation.

Did FB ever take any legal action against Cambridge Analytica? I can't remember anything about it and this sounds very similar to that (although back in those days FB's tools made this incredibly easy).

Re: Taking action against scraping for hire

#68
post #3

Data harvesting is moral for me, but not for thee.

In general I agree that harvesting public data is moral. I think that in these particular cases it's: 1) extracting data from profiles that opted for not being public (only available to logged in users) and 2) reposting scraped data (publicly?) as belonging to the guy who scraped it without users consent.

Facebook has hidden much of Instagram's content behind logins, so that makes most of it "not public".

At the same time, I don't think all of Instagram's users care if their images are hidden, or not.

It's quite unfortunate Facebook/Meta is using hostile language and the word "scraping" together in this case. Scraping is a legitimate process used by various business models to gather information from the Web, which itself was originally intended to be an open forum for people to share content.

Hostile business models have corrupted that intent and turned it into a competitive environment that is harming users and legitimate models which may not have the funding larger corporations can muster.

I have a "scraper" I've built that will either snapshot a page from a user's browser or crawl it remotely with Selinium/Firefox, on the user's behalf, to save the content in an index for searching later, by that user. It's not automated, nor does it parse and crawl URLs in the pages saved. It doesn't use page content in a wider context, either.

I've spent a significant amount of time trying to "work around" anti-scraping efforts by various companies and it's frustrating to see hostility instead of cooperation in certain types of use.

Re: Taking action against scraping for hire

#69
post #29

Earlier quoted context omitted.

Violation of ToS does not mean a violation of the law.

That is why they are suing rather than pressing charges. When someone steals your car you don't sue them you press charges. When someone doesn't uphold their end of a contract you don't press charges you sue for breach of contract.

in reality, you as an individual can't press charges. Only the state can. And many times the state chooses not to. You can sue in civil court, but individuals can't bring cases in criminal court.

Re: Taking action against scraping for hire

#70

Collecting the rhetorical BS: "scraping attacks" Scraping is not an attack. Monopolists want to pretend they own your data because they get unlimited access to monetize it whereas competitors should have none. "self-compromised" Monopolists want to sell you thus it's imperative they maintain the fiction of "one person, one account". By admitting you own your account, they'd have to allow sharing and they wouldn't be…

No post body was provided.
Post reply on HN