Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

941–950 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#941
post #936
post #908

Earlier quoted context omitted.

You're not an author, obviously. Do you put all - and I mean 100% - the code you write in the public domain?

I author a lot of technical documentation actually. Even so, yes. I make all my work public, 100%. I use public repositories as my backup for anything I create, regardless of stage of completion and so others can learn from or help improve my work as they see fit. I do not believe proprietary technology should exist and I put my code where my mouth is. Humans progress faster when we collaborate freely. Fork anything…

So here you say that you put all your work in the public domain, and in a comment above you say that you don't put it in the public domain because you fear a lawsuit?

How the hell does that work?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#942
post #429

Earlier quoted context omitted.

I think the concern goes to the point of copyright to begin with, which is to incentive people to create things. Will the inclusion of copyrighted works in llm training (further) erode that incentive? Maybe, and I think that's a shame if so. But I also don't really think it's the primary threat to the incentive structure in publishing.

> the point of copyright to begin with, which is to incentive people to create things Is it? (I don't agree)

Yes, it is. It's not actually an opinion thing. It's a "what did the people who came up with the idea of copyright think it was for?" thing.

I haven't read the primary source material, so you could teach me something here, but my understanding is that the idea was to incentive creators.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#943

Earlier quoted context omitted.

I think the concern goes to the point of copyright to begin with, which is to incentive people to create things. Will the inclusion of copyrighted works in llm training (further) erode that incentive? Maybe, and I think that's a shame if so. But I also don't really think it's the primary threat to the incentive structure in publishing.

That is not actually the goal of it. Copyright was invented by publishers (the printing guild) to ensure that the capitalists who own the printing presses could profit from artificial monopolies. It decreases the works produced, on purpose, in order to subsidize publishing. If society decides we no longer want to subsidize publishers with artificial monopolies, we should start with legalizing human creativity. Instea…

Interesting! I'd love a citation on this...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#944

Earlier quoted context omitted.

I think the concern goes to the point of copyright to begin with, which is to incentive people to create things. Will the inclusion of copyrighted works in llm training (further) erode that incentive? Maybe, and I think that's a shame if so. But I also don't really think it's the primary threat to the incentive structure in publishing.

People have created for millennia before the modern institution of copyright, so I'm not sure how that's a cogent argument.

Yeah it's an interesting point, but it was also hard to physically copy things for all of those millennia.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#945

Earlier quoted context omitted.

I think the concern goes to the point of copyright to begin with, which is to incentive people to create things. Will the inclusion of copyrighted works in llm training (further) erode that incentive? Maybe, and I think that's a shame if so. But I also don't really think it's the primary threat to the incentive structure in publishing.

i wrote a book and copyright was not once on my mind. having created something is the incentive to create for most artists

I don't think we can infer the motives of most artists from your personal motives.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#946
post #56

Earlier quoted context omitted.

> Google itself got big by indexing other people's data without compensation Weird framing given how much value was and is still placed on Google driving traffic to you

Even before the LLM-craze Google was showing their Answers box or whatever it was called at the top of the results that told you the answer (sometimes) so that you didn’t have to visit any website.

That was significantly after their market dominance was firmly established entirely through sending traffic to external locations.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#947
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

If buying isn't owning, piracy isn't stealing

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#948
I think everyone can see that whatever

(imo not in accordance with the Constitution, after absurdities like deciding “limited time” the way mathematicians might define something of some order of infinity)

the alleged social contract was is not functional the way it was intended, and we see who benefits and who loses.

mass dynamic editing for vitriol and profanity occurred while writing this comment in order to remain within site rules

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#949

Earlier quoted context omitted.

if you want to support an artist go to the show and BUY MERCH at the table! almost all of their income comes from that. the importance of buying a T-shirt at the show cannot be overstated and sometimes you get to say hi to your idol, too

It's a stupid situation, though. There are many creators I'm happy to support - but for 99% of them, I don't want their stupid merch . It's mostly low-quality garbage with high markup, that nevertheless cost something to design and produce, thus wasting both precious resources and labor - an useless tax on contributions to artists that doesn't even help anything. I really wish this wasn't necessary. (Even the okay-qu…

Well, and like the venue will take 50% of merch sales or some , which makes Apple look kinda generous.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#950
post #439

Earlier quoted context omitted.

>What are we actually worried about happening? Few company can amass such quantities of knowledge and leverage it all for their own, very-private profits. This is unprecedented centralization of power, for a very select few. Do we actually want that? If not, why not block this until we're sure this a net positive for most people?

Meta open-sourced it my guy

Because they expect not to have to opens-source future models. Easy to open stuff as long as you strengthen your position and prevent the competition from emerging.

Ask Google about Android and what they now choose to release as part of AOSP vs Play Services.

Post reply on HN