Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

791–800 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#791
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

> the legal system only punishes general public.

In more general terms, the legal system punishes what can be made a profit or an example when punishing.

Also, I don't think the legal system itself wants to get too much into "big institutions against the work of others", save for the fictional TV representations of smart lawyers and clever arguments, 99.9% of the legal system output is copy/paste.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#792

The more I learn about how AI companies trained their models, the more obvious it is that the rest of us are just suckers. We're out here assuming that laws matter, that we should never misrepresent or hide what we're doing for our work, that we should honor our own terms of use and the terms of use of other sites/products, that if we register for a website or piece of content we should always use our work email addr…

> It's only illegal if you get caught

Not quite. It's only illegal if you get caught and you are the wrong kind of person.

For the right kind of person not even a pat on the wrist.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#793
post #685

Earlier quoted context omitted.

> would you rather go back to that today? really?? I absolutely said said no such thing. There are good ways to change things and bad ways to change things. Allowing a private entity reap huge profits by blatantly breaking rules and screwing people is not a good way to change things.

There was no other way to change things on less than a generational timescale. If governments don't like it, well, bummer. They were supposed to serve the people, not the incumbent taxi cartels. They failed, so "we the people" routed around them.

[deleted]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#794
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Google itself got big by indexing other people's data without compensation.

So in other words, it got big by providing free user traffic to people's websites without asking for compensation?

You generally don't charge the phone book money to include you in it. It's actually the other way around.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#795

Earlier quoted context omitted.

His criminality is one matter, but the full weight of the Federal Government on him was an entirely separate matter. A federal prosecutor's job is to jail you regardless of whether it is for downloading a file from a server or for trafficking in humans, and they will come at you with the same vigor regardless of the crime. And nothing has changed about that.

Treating a human trafficker and someone who downloaded some files from a server the same is not in the job description of a prosecutor. What an absurd statement. It's very much the job of a prosecutor to make judgements about the severity of the crime and how to respond. And in this case, the prosecutor showed incredibly poor judgement. There wasn't even a particular reason why the case should go federal in the first…

[deleted]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#796
post #351

Earlier quoted context omitted.

That's what I do, personally. But you call it "stealing," others call it "copying." Stealing takes, from someone, something they own.

There's such a mass of possible works that it hardly constrains someone that if you could cast a magic spell preventing someone from distributing or accessing your particular work and then burned it, your spell would have essentially no effect-- no one would notice it and no one would be harmed. As long as discussion of a work that has published is not impeded, the public is not harmed even by these 50-years after li…

> But it's not a theft of goods, it's theft of service.

What service? If somebody washes your windshield without you asking, it isn't a theft of service to not pay them. A theft of service arises from entering into an agreement and then failing to pay as stipulated in that agreement.

Copyright isn't an agreement you can choose whether to participate in. Copyright is a legal enforcement system that imposes legal liability even on those who don't use it. You may not see this legal liability as "harm", but it absolutely is. Arguing that copyright extends to training is arguing for a dramatic increase in the scope and power of this legal enforcement system.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#797

This reminds me of Peter Sunde's "komimashin" https://www.engadget.com/2015-12-21-peter-sunde-kopimashin.h... It's obviously absurd to enforce copyright as bytes are copied around instead of as it is used. Training an LLM is a different thing than re-hosting and giving away copies to other people. If you don't want people to transform your works - keep them private. You don't own ideas.

Thanks for the link. I wondered what that word meant.

From the article: Kopimashin, as in Copy Machine.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#798

The more I learn about how AI companies trained their models, the more obvious it is that the rest of us are just suckers. We're out here assuming that laws matter, that we should never misrepresent or hide what we're doing for our work, that we should honor our own terms of use and the terms of use of other sites/products, that if we register for a website or piece of content we should always use our work email addr…

If you have a spare few hours, the Acquired podcast episode on Meta is enlightening. They just stumbled through growth hack experiment after experiment without seemingly any risk assessment or ethics.

Excellent podcast, slowly working my way through the back catalogue.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#799

Earlier quoted context omitted.

Copyright and patent aren't the same thing. "Fast moving field" doesn't make sense in terms of copyrights. There's no reason the copywriter should last some minimum duration after the life of the creator. If I write a really popular book, I don't want Hollywood to make it into a movie without compensating me just because they waited a few years

Fast moving field does make sense in terms of copyright because the knowledge is recorded in documents which are then copyrighted. E.g. research papers. > If I write a really popular book, I don't want Hollywood to make it into a movie without compensating me just because they waited a few years I genuinely don't understand this. Even at a decade copyright, pretty much anybody who was going to buy the book and read i…

You know what would happen right?

Copywrite expiring in 20 years doesn't mean access is democratized. Publishers would likely keep the price the same, but instead is the author getting a cut, they just take everything.

Besides. The public isn't owed the fruits of my labor for free.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#800

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

The problem with Uber is that it's a lose-lose-lose money pit scenario.

Customers pay significantly more, drivers make significantly less, and Uber is still running hundreds of millions in the red. Turns out hiring thousands of devs and dumping absurd capital into... driving people around... doesn't really work.

Post reply on HN