Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

501–510 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#501

It really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.

Probably the single biggest thing I learned growing up is that you can safely live by "Everyone is in it for themselves". It's incredibly rare to find people who hold ideals that are detrimental to their own life.

This hasn't been my observation. Instead, I see a society where people regularly help and serve one other, frequently for free. Consider parents, social workers, most academics, food banks, charity in general, most workers in most businesses, et cetera. I wonder: who do you know and work with? A minority of people profit wonderfully off this. Incidentally, they seem to also preach principals that can only lead to the end of their gravy train.

You can counter by insisting that these "altruistic" behaviors are simply less directly but still in the altruist's interest. I would entirely agree.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#502
post #459

Earlier quoted context omitted.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

It is possible that digitization and improvement of taxi services was inevitable anyway

Not only was it inevitable, if we were so inclined and willing to use the regulatory pen, we could've simply written into law that for Taxi's to operate, they must be well maintained and must accept all major forms of payment. And yeah, the Taxi industry would've fought it because every company ever has fought every regulation ever no matter how much it stands to benefit both their customers and they them-fucking-selves but companies having a say in how they are regulated is both how a Taxi company would fight this, and how Uber, AirBnb, OpenAI, Meta, etc. blatantly and flagrantly violate the law and instead of consequences, they get fines, and court hearings. So maybe we just shouldn't be allowing that?

It drives me up the goddamn wall how people will say shit like "the Taxi industry needed to be upended" when like... I mean, maybe? But on balance, given all the negative externalities associated with these companies, are they really a gain? Or are they just a different set of overlords, equally disinterested in providing a good service once they reach the scale where they no longer are required to give a shit?

Just... regulate the fuckers. Are you sick of filthy Taxis that break down? Put a regulation down that says if a cab breaks down during a trip, they owe the customer a free ride and five thousand dollars. You bet your ASS those cabs will be serviced as soon as humanly possible. This isn't rocket science y'all. Make whatever consequence the government is going to dispense immeasurably, clearly worse than whatever the business is trying to weasel out of doing, and boom. Solved.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#503
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

When individuals are assigned heroic status despite clear evidence of mental illness and crimes, such as “breaking and entering”, it prevents society from having rational discussions about both law enforcement and mental health support. This dynamic repeats across multiple high-profile cases.

People often elevate deeply flawed figures to heroic status when those figures seem to challenge authority or "the system." This happens especially with individuals who present themselves as outsiders fighting the establishment, have a compelling personal struggle narrative, or voice grievances that resonate with public frustrations

Trump fits this pattern - his supporters overlook concerning behaviors and statements because they see him as fighting a system they distrust. Like Manning and Swartz, his mental state and fitness are often ignored in favor of the "hero against the system" narrative.

This dynamic creates a feedback loop where legitimate criticism becomes harder to discuss rationally.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#504
post #422
post #339

Earlier quoted context omitted.

I definitely do know what I think.

Do you think developing countries are just peppered with libraries, and their inhabitants order books from Amazon? Libgen originated in Russia, and its users are global. This is not a purely American issue.

I was responding to a comment arguing that LibGen is the largest collection of knowledge in human history, which I think is an overly romantic and totally incorrect take. It may be a very useful collection of knowledge to people in developing countries, but it simply is not larger than the collection of knowledge accessible via any first-world public library. Obviously not everyone has access to that, but again, that’s not what the comment I replied to was about, and not relevant to what I’m saying.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#505
LLMs are worse than search for figuring out what value a specific asset provides to the LLM. Atleast with search your work or page is not lost and still gets a click/user interaction, and may be give you a chance to monetize the interaction. However, LLMs just don’t have any such option. Gemini adds links but the links they add are completely editorialized by the LLM and need not reflect the original at all. So how does anyone ask for compensation even if they sue?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#506

Earlier quoted context omitted.

It's a stupid situation, though. There are many creators I'm happy to support - but for 99% of them, I don't want their stupid merch . It's mostly low-quality garbage with high markup, that nevertheless cost something to design and produce, thus wasting both precious resources and labor - an useless tax on contributions to artists that doesn't even help anything. I really wish this wasn't necessary. (Even the okay-qu…

What is the point of this comment? Just a stream of consciousness for a future LLM sweep? Nobody thinks that the actual creator should get nothing. Are you asking for better T shirts? Do you want more direct ways of just giving cash?

I agree mostly with their comment.

What I want more than anything is for bands to just sell me a damned CD. I've lost track of how many times an artist I want to support doesn't release their music on CD. I'd even settle for DRM-free flacs, if it costs less than a CD.

High quality sheet music would be cool. Lindsey Stirling is the only artist I can think of that does that though. Rasputina used to at one point.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#507
post #156

Earlier quoted context omitted.

What happens in US right now shows that change is achieved through voting. There are other examples as well in Europe where things did change because of how people voted. If the change is good or bad depends on your perspective. For me the annoying part is that people vote for a guy because of a couple heavily advertised issues, ignoring all the other plans or the fact that he might not keep his word. Then they are u…

I like your optimistic take. My more cynical one is that what’s happening in the US shows that real change is achieved through corruption and lying: honest policy discussions and iterative improvement stand no chance against a charismatic populist who will say anything to entrench an oligarchy.

It's not primarily optimistic. I just think that education of the people can bring the best improvement on the long run, and not adjusting democracy, demonizing rich guys or another "new" system.

While I hope iterative improvement is the way, I think there are people that have it (or feel) so bad (due to various reasons) that they would take a 50% chance to die for the chance to live better.

The charismatic populists are not supported only by people that are well off, without any worry (neither in US, nor in Europe). (ex: https://www.statista.com/statistics/1535295/presidential-ele...)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#508

Earlier quoted context omitted.

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

It's an oligarchy, always has been. I don't know how colossal the pile of evidence supporting this has to get before people finally accept it.

Conscious life in general seems possible to me unless our brain tells us a better story than reality.

A story in which we are the hero, in which we are not mortal, in which we are important, in which people care about us, in which we are intelligent and our perceptions rarely fail us, in which our life has a meaning and also in which the social game we play is determined, or at least influenced, by some just principles. We would despair if we were aware of the full extent of our meaninglessness and powerlessness.

I believe that it is the core reason why we love to believe that God/Nature is good, that the king is legitimate and that the laws are fair.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#509
The question is, if they could and would have paid for each book, would it be ok to train the LLM on them? I'm talking about prior books, I'm sure new books have language forbidding their use to train LLMs at the point of sale. But legally, how does using a book to train a LLM differ from a teacher learning from a book and teaching its contents to their pupils. Obviously, the LLM can do so at scale, but is there a legal difference?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#510
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

> MIT

I think Aaron Swartz went to Harvard, not MIT

https://en.wikipedia.org/wiki/United_States_v._Swartz

Post reply on HN