Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

551–560 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#551
post #436
post #391

Earlier quoted context omitted.

> Are AI-written books getting published? actually i think they are. lots of e-book slop > If they start out-competing humans, is that bad? Not inherently, but it depends on what you mean by out-competing. Social media outcompeted books and now everyone's addicted and mental illness is more rampant than ever. IMO, a net negative for society. AI books may very well win out through sheer spam but is that good for us?

Nobody has responded to me with anything about how authors are harmed, so I don't really get who we're protecting here. It feels more like we just want to punish people, particularly rich people, particularly if they get away with stuff we're afraid to try.

> Nobody has responded to me with anything about how authors are harmed

i imagine if books can be published to some e-book provider through an API to extract a few dollars per book generated (mulitiplied by hundreds), then eventually it'll be borderline impossible to discover an actual author's book. breaking through for newbie writers will be even harder because of all of the noise. it'll be up to providers like Amazon to limit it, but then we're then reliant on the benevolence of a corporation and most act in self interest, and if that means AI slop pervading every corner of the e-book market, then that's what we'll have.

kind of reminds me of solana memecoins and how there are hundreds generated everyday because it's a simple script to launch one. memecoins/slop has certainly lowered the trust in crypto. can definitely draw some parallels here.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#552

Earlier quoted context omitted.

[flagged]

I don't think it's right to downplay the disproportionate response the FBI had to Aaron's actions. He was initially being threatened with 50 years in prison and a $1 million fine, the stress from which sent his mental health spiraling and in no small way contributed to his suicide. I think the original point of the person you are responding to still stands.

You are correct. Prosecutors don’t consider whether charging a perp will upset them and stress them out. Nor should they.

Don’t do the crime if you can’t do the time.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#553

Earlier quoted context omitted.

[flagged]

For downloading papers , paid for with government funding and gatekept by greedy rent seekers charging ~ $30 a pop. The lengths people will go to defend things that should not exist astounds.

What about his family and friends? Do you blame them as well?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#554
post #529

Earlier quoted context omitted.

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

If you do something wrong then you, as a person, are held responsible and accountable. If you do something wrong as "part of your job" then you're typically not held responsible and accountable but the company is (the exceptions being spectacular fraud: Enron, VW diesel). It's not hard to see how this can go off the rails.

“The revolution will be incorporated.”

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#555
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

I think if Google attempted to download the entirety of JSTOR with the express intent of making the full dataset freely available, then Google would also face legal consequences. It's true, and relevant, that Google would feel those consequences much less sharply than Swartz did.

Google Scholar explicitly made direct deals with publishers to scrape their content, with the constraint that while they can use the content to serve search results in Scholar, but cannot show the content of the papers on the site- just titles and short fragments that match. the deals were tenuous and I had to step carefully around my plan to use that database to implement large-scale scientific search over the literature (this was a long time before anybody was seriously considering using LLMs on research data).

I've spoken to several very wealthy/powerful people and tried to get them to negotiate a large-scale content license with the various publishers that would allow researchers and individuals to access more research in lower-friction ways. None of them (NIH, Schmidt, etc) were really interested.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#556
post #166

Earlier quoted context omitted.

Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

Part of the reason Airbnb got a pass must be how profitable it was to people who own many properties, despite the harm it does to the communities of people who only own one property.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#557
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

And Hollywood was created on the west coast because for intellectual property it was still the far west and it allowed them to ignore patents on movie technologies.

They became the thing they lamented.

This is the inevitable.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#558
post #515

Earlier quoted context omitted.

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life. In case anybody here doesn't know, that's a reference to Aaron Swartz, an activist (and Reddit co-founder) that was risking 35 years in prison and a $1 million fine just for downloading a lot of academic papers from JSTOR. He eventually took his life because of the pressure. May his soul rest in peace.

Except he was offered 6 months in a plea bargain, which he declined because he wanted a trial. Whether 6 months was reasonable punishment for "plug a laptop into a closet at MIT to download some scientific papers" is another matter, but "you forfeit your life" or "35 years in prison and a $1 million fine " is massively misleading.

[dead]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#559

Earlier quoted context omitted.

They could have only leached and refrained from sharing any part of copyrighted data. If i were to commit something as risky as this, that is what i would do.

Then it would need to be determined, whether that is the case or not. Did every single machine they used have the configuration for only leeching and no seeding? The company is liable for what its employees on the job. If only one employee was also seeding ... that could be a very interesting case.

> Did every single machine they used have the configuration for only leeching and no seeding?

I would certainly assume so. It's incredibly obvious that's what you would want to do from a legal standpoint.

> If only one employee was also seeding ... that could be a very interesting case.

The torrenting wouldn't be done casually by employees acting on their own. And it's not like multiple employees are doing it simultaneously, unsupervised, on their personal computers.

This is part of an official project. They'd spin up a machine just to download the torrent, being careful to disable seeding.

This is Meta. They have lawyers involved and advising. This isn't a teenager who doesn't fully understand how torrenting works.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#560
post #246

How about a consequentialist argument? In some fields, AI has already surpassed physicians in diagnosing illnesses. If breaking copyright laws allows AI to access and learn from a broader range of data, it could lead to earlier and more accurate diagnoses, saving lives. In this case, the ethical imperative to preserve human life outweighs the rigid enforcement of copyright laws.

There’s nothing particular to AI about your comment, it’s a general downside of IP.

No, the development of an artificial general intelligence does seem like a special case compared to usual IP debates, particularly in the potential multiplicative positive-sum effects on society overall.
Post reply on HN