Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

541–550 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#542

Earlier quoted context omitted.

Meta argues that it's fair use, and that they just downloaded, and never seeded, all the torrents.

Seeding and downloading are in the same protocol. You can't do one without the other

Why comment if you have no idea what youre talking about?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#543

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

It really screwed over a lot of regular working-class people. In some European cities getting a taxi license was a serious monetary investment. People took our huge loans for this. This was now suddenly worthless. It's like being told your very expensive university education is no longer accredited, but the student loan still exists. kthxbye.

I'm not saying the existing systems were always good (they weren't), but you need to be willing to overlook a lot of real-world suffering to be "rooting for Uber". Phrases like "taxi cartels" sound nice, but they're hardly neutral phrasings that simplify things to the point of being useless phrases.

And "I'm just going to willingly and knowingly ignore laws I don't like for personal profit" is not a great take-away either. This isn't Aaron Swartz breaking a law as a matter of "civil disobedience" – it's just a plain "how can we make money?"

And where does that leave competitors who are NOT willing to break the law? It's an unlevel playing field; there can be no free market if some people don't need to follow the same set of rules. Uber's actions are fundamentally anti-capitalist and anti-free market.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#544

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

Uber definitely improved things.

When traveling it’s also so much safer than taxis.

My brother was robbed at gunpoint in a taxi. My wife had to jump from more than one moving taxi to escape. My ex girlfriend too. My Swiss friend had his camera and wallet stolen.

You can have issues with Uber too, but not as frequently because there’s a digital audit trail, you can report them to the platform and the police. The threat of those consequences lead to better behavior.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#545
At OpenAI we have seen some employees expressed their concern publicy about the moral grounds on which company was acting. We never heard about it from anyone at Meta but there were some jokes ofcourse. I guess everything is fair in AI and Corporates.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#546

The question is, if they could and would have paid for each book, would it be ok to train the LLM on them? I'm talking about prior books, I'm sure new books have language forbidding their use to train LLMs at the point of sale. But legally, how does using a book to train a LLM differ from a teacher learning from a book and teaching its contents to their pupils. Obviously, the LLM can do so at scale, but is there a le…

> The question is, if they could and would have paid for each book, would it be ok to train the LLM on them?

Whether training on AI model on an array of diffentent works, many of which are copyright protected, is itself a copyright violation, in addition to or distinct from any copyright violation that goes on gathering the dataset for training (and separate from any copyright violation in the actual or intended use of the LLM), remains to be resolved as a legal question, and may or may not have a simple yes or no answer (or the same answer under every system of copyright laws globally).

My inclination is that it is probably generally not a violation in US law, but that's not something I am very confident in; how the definitions of copy and derivative work apply to determine if it would be without fair use, and how fair use analysis applies, are not clear from the available precedent.

> But legally, how does using a book to train a LLM differ from a teacher learning from a book and teaching its contents to their pupils.

It is very clear, by looking at how US copyright law is written and even more clear in its history of application, that information stored in brains of people are without exception neither copies nor new works that can be derivative works under US law, and so cannot be infringing, no matter how you gain them. It’s also very clear in the statute itself and the case law that data in media used by artificial digital computers, on the other hand, can constitute copies or derivative works that can be infringing. Even if the process is arguably similar in legally relevant manners, copyright law is critically focussed on the result and whether it is a particular kind of thing which can be infringing, not just the process.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#548

Earlier quoted context omitted.

[flagged]

Everyone is responsible for the full effects of his actions. One is literally responsible for all consequences, everything no matter how indirect. This absolute responsibility is physics, while the limited 'only direct consequences' type thing is a choice made in some human legal systems. People are smart. They know what stress they put on people and from interacting with them they get a good feel of much they can ta…

[deleted]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#549

Earlier quoted context omitted.

In Spotify’s defense, they used the pirated data only to show a proof of concept to the copyright holders, and that use was sanctioned by the local rights holders organization STIM. The copyright holders then approved their concept, and subsequently Spotify got the rights to offer their service to customers. Everybody won.

That’s not entirely true, in Spotify’s early days you could upload files to the service and listen to songs uploaded by other people. I think the majority of any song I wanted to listen to before they went Europe-only for a time was “pirated”.

Indeed. I remember there was one song that used a pirated variant and you could tell because it had an obvious artifact that was accidentally introduced in a pirated copy of the song!

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#550
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Yes. And the problem here isn't that companies get away with doing things like this, the problem is that individuals don't. Attempting to lock information behind a nightmarish legal system is the problem. I'm pretty much at the point now where I don't buy the "copyright incentivizes creation" argument any more. Copyright, like advertising, incentivizes creation by enormous corporations, but also like advertising it i…

Nothing stops you from downloading Ann’s archive and training a model on it, right? The likelihood that you, as an individual, get sued over is is virtually zero.

This is what Meta tried to do, quietly download and use the data, to do research and advance their LLMs, without trying to establish any legal precedents or pick up fights.

Post reply on HN