Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

241–250 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#241
post #209

Libgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.

Libgen turns into a problem when you have a company developing generative AI with it, either giving money to GPU manufacturers or themselves with paid services (see OpenAI)

…why? Will people buy less books because we have intuitive algorithms trained on old books?

Personally, I strongly believe that the aesthetic skills of humanity are one of our most advanced faculties — we are nowhere close to replacing them with fully-automated output, AGI or no.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#242
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Spotify's music library was also pirated in the early days. I want to know more, please enlighten me (anyone who knows). I read the book "The Spotify Play" and it made it seem like the pirated music was an internal-only thing and not something available to customers. Is that true?

Before the launch, Spotify had a deal with the music rights holders association in Sweden (STIM) that they could use a merged collection of friends and families music libraries. All this was removed before Spotify went out of beta.

So while it was using pirated media, it was sanctioned by the rights holders for the experiment of building Spotify.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#243
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

[flagged]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#244

Earlier quoted context omitted.

the english empire once tried to mantain a monopoly over steam loom machines the americans cheated their way to competition, heck, even before that, the english empire got jumpstarted by stealing gold from the spanish (who were themselves exploiting it away from aztec and other mexican natives) I'm saying it's business as usual, but also, culture doesn't work like tangible physical widgets so we must stop letting a f…

I don't think I've heard the term "English empire". Is it an attempt by the Scottish to pretend they weren't involved?

I was assuming they were talking about pre-1706 given the Spanish gold context.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#245
Yes it smells bad but facebook did the right thing (at least for facebook)

After OpenAI trained their models on the famed books2 dataset, and seeing the technological implications of ChatGPT, there was a good chance they would let them get away with it.

Would the USA really surrender its AI technological advantage for trivial matters like copyright? They would make some royalty arrangement and get it over with

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#246
How about a consequentialist argument? In some fields, AI has already surpassed physicians in diagnosing illnesses. If breaking copyright laws allows AI to access and learn from a broader range of data, it could lead to earlier and more accurate diagnoses, saving lives. In this case, the ethical imperative to preserve human life outweighs the rigid enforcement of copyright laws.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#247

I strongly urge people to read Thomas Babington Macaulay's speeches on copyright, its aims, terms, and hazards. Very well reasoned and explained. In particular, people often cited the case of authors who had died leaving a family in destitution, and claimed that copyright extension would be a fair way of preventing this, but in most cases the remaining family had never held the copyright; the author had initally sold…

> in most cases the remaining family had never held the copyright; the author had initally sold the reproduction rights to a publisher

He was able to sell it because it is something valuable, exactly because of the copyright protections. Regardless of whether author sells the rights or not, he and his family would equally be better off with copyright.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#248
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

"Zuckerberg was at White House for meetings on Thursday" - https://www.reuters.com/world/us/zuckerberg-was-white-house-...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#249

I strongly urge people to read Thomas Babington Macaulay's speeches on copyright, its aims, terms, and hazards. Very well reasoned and explained. In particular, people often cited the case of authors who had died leaving a family in destitution, and claimed that copyright extension would be a fair way of preventing this, but in most cases the remaining family had never held the copyright; the author had initally sold…

> Thomas Babington Macaulay The one who got Hindu Sanskrit books translated in a horrible manner and then claimed: "I have no knowledge of either Sanskrit or Arabic. But I have done what I could to form a correct estimate of their value. I have read translations of the most celebrated Arabic and Sanskrit works. I have conversed both here and at home with men distinguished by their proficiency in the Eastern tongues.…

This is the corollary of the fallacy of appeal to authority: the rejection of an argument on the grounds that the speaker was horribly wrong on an unrelated or very loosely related topic.

If you reject Macaulay on copyright because he was an imperialist, you can use the exact same logic to reject the arguments of essentially every person who ever lived. Very few humans who ever wrote anything important will perfectly align with your morality, and most will be horribly misaligned in at least one way.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#250
post #156

Earlier quoted context omitted.

It doesn't "seem". The entire system in most countries works, by design, that way because the people in power trade in influence at a different plane. That's why democracy often feels "failed" in that no change can be achieved because "it's just more of the same". Few Lobbyists representing the interests of a few people have more power than millions voting differently.

What happens in US right now shows that change is achieved through voting. There are other examples as well in Europe where things did change because of how people voted. If the change is good or bad depends on your perspective. For me the annoying part is that people vote for a guy because of a couple heavily advertised issues, ignoring all the other plans or the fact that he might not keep his word. Then they are u…

I like your optimistic take. My more cynical one is that what’s happening in the US shows that real change is achieved through corruption and lying: honest policy discussions and iterative improvement stand no chance against a charismatic populist who will say anything to entrench an oligarchy.
Post reply on HN