Live data from Hacker News

Taking action against scraping for hire

about.fb.com

141–150 of 240 posts

Re: Taking action against scraping for hire

#141
Funny story from the early days of TheFaceBook, probably around 2005ish:

I was a webmaster of a set of servers on a major university's network. I also had access (enough to run arbitrary programs that had pretty much full ingress/egress to the public internet) to a number of machines across the campus's network. Through some of my coursework and ACM chapter activities I met some other similarly minded technical people with similar levels of access.

We decide that it would be fun to use our superpowers (access + programming abilities + curiosity) to sign up for various accounts on FB and essentially scrape and friend as much as possible. At the time they had some rate limiting, some IP banning (which wasn't terrible because the Uni gave public IPv4 addrs to all machines on campus by default) and then added some early CAPTCHA which we ended up breaking pretty trivially with some python and image recognition code.

Never got sued... :) Never really did much with the scripts or data except test that they worked. Fun times.

Re: Taking action against scraping for hire

#142

Collecting the rhetorical BS: "scraping attacks" Scraping is not an attack. Monopolists want to pretend they own your data because they get unlimited access to monetize it whereas competitors should have none. "self-compromised" Monopolists want to sell you thus it's imperative they maintain the fiction of "one person, one account". By admitting you own your account, they'd have to allow sharing and they wouldn't be…

Quoted post unavailable.

> the vast majority of Web scraping efforts are to build businesses on top of other organizations hard work and innovation. Period. End of story.

Yeah and the vast majority of the internet and all these mega corps run on open source while paying pittance back to the ecosystem. Cry me a fuckin river.

Can't wait til someone sue's them for "scraping" their site for web previews and thumbnails everytime someone shares a link on Facebook.

The double standard of these muppets.

Re: Taking action against scraping for hire

#143
post #57

Earlier quoted context omitted.

This is a false expectation and it’s important people learn this.

They’ll stop posting in the way they currently enjoy and will, therefore, have lost some freedom. Great outcome! In other news: your partner may also leak your most intimate secrets. I hope they do, to teach you a lesson? Every trust can be betrayed. Why do you believe a world without trust would be better? Only because you cannot handle the nuance of different levels of trust?

So taking shackles off is called “losing freedom” now? Also, people enjoy many things, just look at the junkheads. Still, it's more natural to have trust in a heroin addict than to have trust in businesses like Facebook.

Re: Taking action against scraping for hire

#144
post #85

Earlier quoted context omitted.

> the vast majority of Web scraping efforts are to build businesses on top of other organizations hard work and innovation. Not really. Scraping just gets data, not code, so it's hard to support this argument. The anti-scraping view is that the right to use the data rests with the company that collected it, but I don't think that view is held by most people.

If you are arguing that an organization's data is worthless but only their code has worth, then I'm not quite sure where to go from this point in this discussion, other than to say that is crazy .

The data is obviously valuable, but they don't necessarily deserve a monopoly on that data, since that data primarily belongs to the users who created the data; so while it's understandable that organizations want to restrict that data, we have no obligation (moral or otherwise) to respect that desire.

Re: Taking action against scraping for hire

#145

Collecting the rhetorical BS: "scraping attacks" Scraping is not an attack. Monopolists want to pretend they own your data because they get unlimited access to monetize it whereas competitors should have none. "self-compromised" Monopolists want to sell you thus it's imperative they maintain the fiction of "one person, one account". By admitting you own your account, they'd have to allow sharing and they wouldn't be…

Quoted post unavailable.

In my opinion, breaking a click-through license agreement or violating the small print on some dense and difficult to read web page is hardly an issue of morality or ethics.

Let's also remember that a big reason Meta is hating on scraping is because of their own problematic behavior. It wasn't so long ago that they were suing NYU over research on political ads and how Facebook targets their readers.[0] In fact, it wouldn't surprise me if Meta's larger goal is to prevent this sort of research.

[0]: https://news.bloomberglaw.com/privacy-and-data-security/face...

Re: Taking action against scraping for hire

#147

Collecting the rhetorical BS: "scraping attacks" Scraping is not an attack. Monopolists want to pretend they own your data because they get unlimited access to monetize it whereas competitors should have none. "self-compromised" Monopolists want to sell you thus it's imperative they maintain the fiction of "one person, one account". By admitting you own your account, they'd have to allow sharing and they wouldn't be…

Quoted post unavailable.

I feel the same way. My biggest pet peeve is that scrapers/bots traversing my site generates more data than the target audience of users. The scrapers get all of this data for "free" at my expense of the hosting costs to provide them that "free" data.

Re: Taking action against scraping for hire

#148
post #102

Earlier quoted context omitted.

It's not, see https://www.urbandictionary.com/define.php?term=Simp

I can also link to a source that's going to be biased in my favor: https://www.etymonline.com/word/simp

> 1903

> 1640s

Lol, no. I'm using the definition from this century:

> Someone who does way too much for a person they like

Re: Taking action against scraping for hire

#149
post #44

Earlier quoted context omitted.

I don't think I know the answer, but I'm curious: Does violating a website's TOS meant your accessing it beyond your authority, making it a violation of the US's Computer Fraud and Abuse Act?

Violating TOS no; Gaining access beyond your authority maybe https://www.eff.org/deeplinks/2010/07/court-violating-terms-...

I was assuming that in this case, a person's authority was specifically granted by the ToS.

I wondered if the interplay of those two concepts muddied the waters.

Re: Taking action against scraping for hire

#150

Collecting the rhetorical BS: "scraping attacks" Scraping is not an attack. Monopolists want to pretend they own your data because they get unlimited access to monetize it whereas competitors should have none. "self-compromised" Monopolists want to sell you thus it's imperative they maintain the fiction of "one person, one account". By admitting you own your account, they'd have to allow sharing and they wouldn't be…

The only argument I have here (sadly in favor of FB) is with "safeguard people against clone sites". While I did give my data to FB, I didn't approve that transfer to another site/system. That is the only place I could possibly see some legal foot hold.
Post reply on HN