Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…
Meta torrented & seeded 81.7 TB dataset containing copyrighted data
301–310 of 981 posts
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#302Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#303Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#304Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#305I'm pretty sure you can theoretically download torrents without seeding, although this is frowned upon. If they really seeded (with full bandwidth?) that's indeed pretty brazen.
It is sort of strange that Meta is being singled out here though, and sort of sad considering they at least release the model weights. What's the signal? Do illegal shit to be competitive, but make sure there is no evidence?
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#306Earlier quoted context omitted.
Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.
The hotel and taxi industry were legit terrible before those two disrupted them. Laws are ment to be broken. Especially in cronist systems where incumbents write the laws.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#307Earlier quoted context omitted.
> Thomas Babington Macaulay The one who got Hindu Sanskrit books translated in a horrible manner and then claimed: "I have no knowledge of either Sanskrit or Arabic. But I have done what I could to form a correct estimate of their value. I have read translations of the most celebrated Arabic and Sanskrit works. I have conversed both here and at home with men distinguished by their proficiency in the Eastern tongues.…
This is the corollary of the fallacy of appeal to authority: the rejection of an argument on the grounds that the speaker was horribly wrong on an unrelated or very loosely related topic. If you reject Macaulay on copyright because he was an imperialist, you can use the exact same logic to reject the arguments of essentially every person who ever lived. Very few humans who ever wrote anything important will perfectly…
On the contrary I would argue that this is precisely why you SHOULD NOT take his opinion on copyright. One of the main outcomes of imperialism/colonization is denigrating/destroying/appropriating works of art, literature with the primary goal of subjugation, subversion and thereupon replacement of native culture/traditions/institutions. I did not quote the other half of his nauseating take but I'll post it nevertheless:
"[...] And I certainly never met with any Orientalist who ventured to maintain that the Arabic and Sanscrit poetry could be compared to that of the great European nations. But when we pass from works of imagination to works in which facts are recorded, and general principles investigated, the superiority of the Europeans becomes absolutely immeasurable. It is, I believe, no exaggeration to say, that all the historical information which has been collected from all the books written in the Sanscrit language is less valuable than what may be found in the most paltry abridgments used at preparatory schools in England. In every branch of physical or moral philosophy, the relative position of the two nations is nearly the same."
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#308Earlier quoted context omitted.
The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.
The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#309I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3…
1.) Training on copyright that is publicly available. You write a poem and publish it online for the world to read. That is your IP, no one else can take it an sell it, but they are free to read and be inspired by it. The legalitly of training on this is in the courts, but so far seems to be going in favor of LLMs.
2.) Training on copyright that is not publicly available. These are pretty much pirated works or works obtained by backdoor to avoid paying for them. Your poem is behind a paywall and you never got paid, yet the poem is known by the LLM. This is just straight illegal, as you legally must pay to view the work. However there might be conditions here too like paying for access to an archive and then training on everything in it.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#310I strongly urge people to read Thomas Babington Macaulay's speeches on copyright, its aims, terms, and hazards. Very well reasoned and explained. In particular, people often cited the case of authors who had died leaving a family in destitution, and claimed that copyright extension would be a fair way of preventing this, but in most cases the remaining family had never held the copyright; the author had initally sold…
I’m a huge IP hater and am sure that happens, but to be fair, letting copyright extend past death also increases the amount the author can sell it for in the first place.