Live data from Hacker News

Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

blog.curiousquail.com

411–420 of 427 posts

Re: Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

#411
post #364
post #329

Earlier quoted context omitted.

Amusing, but a long string of rogue cyber intrusions coupled with Jstor blocking MIT altogether a few times isn’t something to ignore. Hard to imagine just ignoring the risks and potential of that unknown. MIT and Jstor had a bunch of incident response going on, and it isn’t a leap to consider Jstor would cut them off for good if the situation went unresolved.

Why do you seem so desperate to characterize Aaron as some amoral super-spy hacker? His ethics are well known. He took issue with Jstor's business model of gatekeeping tax-payer-funded research behind another pay wall, stole some articles, got caught, settled with jstor and MIT. Life could have gone on. The OP's blog post is about how the law isn't applied equally in the United States. Aaron was an individual and an…

There are two entirely different issues:

1. Is the law aligned with moral and ethical expectations? Probably not.

2. Is the process reliable? At least since the Derek Chauvin trials, I'm having doubts, but it doesn't seem it had failed in this case.

Sure, cases of the former need urgent fixing (and we're not getting that), but the latter scenario falls into the "The end is nigh" category.

Re: Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

#412
post #400

Earlier quoted context omitted.

Both of your supposed counterexamples are Chinese. China does things differently from the USA.

My counterexamples prove exactly the point; TBTF is fugazi, arbitrary, something in your head.

Because two companies in China aren't it, it doesn't exist?

Re: Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

#413
post #408

Earlier quoted context omitted.

Anyone can scan, host, and curate - including one or many peers. We do not need any entity in between researchers and other researchers.

Practically speaking, no-one is going to scan centuries' worth of historical journal articles for free. And JSTOR isn’t really an entity outside the research community; it’s a non-profit that grew from within it.

I think we’re talking past each other mostly because I start to lose value on most research for my topics of interest that are more than two or three decades old, let alone five or six.

This may be different for others.

Re: Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

#414
post #408

Earlier quoted context omitted.

Practically speaking, no-one is going to scan centuries' worth of historical journal articles for free. And JSTOR isn’t really an entity outside the research community; it’s a non-profit that grew from within it.

I think we’re talking past each other mostly because I start to lose value on most research for my topics of interest that are more than two or three decades old, let alone five or six. This may be different for others.

So, this discussion started with me saying that:

> [JSTOR is] a non-profit that scans old journal articles that would otherwise be a huge pain to access

You'd think that would pretty obviously include the vast swathes of material published before the internet and PDF preprints were a thing (and the considerable amount of subsequent research that just never got uploaded to anyone's website, for whatever reason).

Re: Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

#415
post #414

Earlier quoted context omitted.

I think we’re talking past each other mostly because I start to lose value on most research for my topics of interest that are more than two or three decades old, let alone five or six. This may be different for others.

So, this discussion started with me saying that: > [JSTOR is] a non-profit that scans old journal articles that would otherwise be a huge pain to access You'd think that would pretty obviously include the vast swathes of material published before the internet and PDF preprints were a thing (and the considerable amount of subsequent research that just never got uploaded to anyone's website, for whatever reason).

Okay, so why can’t the volunteers scanning for Anna’s Archive perform this again?

Why does it need to be technically centralized and politically weak?

Re: Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

#416
post #414

Earlier quoted context omitted.

So, this discussion started with me saying that: > [JSTOR is] a non-profit that scans old journal articles that would otherwise be a huge pain to access You'd think that would pretty obviously include the vast swathes of material published before the internet and PDF preprints were a thing (and the considerable amount of subsequent research that just never got uploaded to anyone's website, for whatever reason).

Okay, so why can’t the volunteers scanning for Anna’s Archive perform this again? Why does it need to be technically centralized and politically weak?

Anna's Archive is primarily a meta search engine. I see that they've put out a call for volunteers, but I don't think any significant part of their archive consists of volunteer scans.

As to your 'why' question, we're talking about boring, thankless work that in many cases breaks the law. It's not exactly surprising that we don't see millions of people signing up to do it for free!

I don't understand your second paragraph.

Re: Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

#417

Earlier quoted context omitted.

The difference here is that the scientific papers he downloaded weren't freely available to the public, like those scraped webpages would be. Corporate scrapers have been sued[0] in the past for scraping pages from behind a login page / paywall. [0]: https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn

Neither are the pirated books Meta is using for model training.

I totally agree with you, and didn't mean to endorse that.

Re: Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

#418

Earlier quoted context omitted.

>The public might donate food to the poor. That doesn't give them the right to go into their house and rummage through their fridge. I think a better analogy for this situation is: The public donates food, then the recipient, after being fed, sells access to (infinite cheaply replicable copies of) said food for a profit. The public in this case just wants to have said food.

Maybe we are better off avoiding the analogies. The public gives them cash for terms of a grant. If those dont include an open access paper, it is unreasonable to demand it after the fact. It certainly doesn't give a right to go take take their papers (or whatever they made).

There's some common-sense understanding of this situation that I'm not sure you're wilfully overlooking or oblivious to. The claim is that it's unjust to blatantly exploit a system that was intended to fund/sustain you, for profit. It's not necessarily illegal behavior since "open access" was not in the terms, but how can it possibly be morally defensible that they have managed to find a way to legally subvert/trick the system for their own benefit? An overly pedantic view of the situation isn't helpful.

Re: Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

#419
post #416

Earlier quoted context omitted.

Okay, so why can’t the volunteers scanning for Anna’s Archive perform this again? Why does it need to be technically centralized and politically weak?

Anna's Archive is primarily a meta search engine. I see that they've put out a call for volunteers, but I don't think any significant part of their archive consists of volunteer scans. As to your 'why' question, we're talking about boring, thankless work that in many cases breaks the law. It's not exactly surprising that we don't see millions of people signing up to do it for free! I don't understand your second para…

I simply don’t see JSTOR outlasting (as an organization, or a technology/archive) individual contributors and torrent trackers.

Solution which does not require millions of people =]

Re: Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

#420
post #416

Earlier quoted context omitted.

Anna's Archive is primarily a meta search engine. I see that they've put out a call for volunteers, but I don't think any significant part of their archive consists of volunteer scans. As to your 'why' question, we're talking about boring, thankless work that in many cases breaks the law. It's not exactly surprising that we don't see millions of people signing up to do it for free! I don't understand your second para…

I simply don’t see JSTOR outlasting (as an organization, or a technology/archive) individual contributors and torrent trackers. Solution which does not require millions of people =]

At present I don’t see individual contributors making any significant contribution towards scanning old academic journals. Why would this change in the future? There was at least a decade where the technology to enable this existed, and where most older issues of most journals were not available online, and yet the army of volunteer scanners failed to materialize. Now that there is already a non-profit doing this archival work at very high quality and with economies of scale, it seems even less likely that such a volunteer effort will materialize. But if you want to prove me wrong, get off HN and go scan some old journal articles!
Post reply on HN