Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

491–500 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#491
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

How does that prosecutor sleep at night?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#492
post #329

Earlier quoted context omitted.

I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.

They have no morals, therefore I shouldn't either! That'll teach 'em!

The observation being made here is that copyright law serves to protect the interests of large companies, not the public, so violating copyright law is, in and of itself, not unethical.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#493

Earlier quoted context omitted.

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life. Just to point, but the material in question was public domain, so nobody had even a copyrights claim over it.

Do you have a citation for that claim? I've not seen a claim that none of the material had copyright before.

It's a library of historical scientific work. You will find the famous Einstein's 3 1905 papers there, for example.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#494
post #124

Earlier quoted context omitted.

Too much paperwork, too much effort. These are important people, doing much more important stuff than whatever book authors do. Or so they think, I think.

I doubt they think that way, but even if they did, they'd be right - for 99% of the works in question, the biggest value they gave to the world is, by far, being part of the LLM training corpus. There's lots of content out there. Most of it is noise. People forget because they're only ever exposed to an aggressively curated fraction of it.

No, it's not.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#495
Damn! One of my old books can be found in the Anna's Archive search. The book has been out of print for years. I pity the Meta users who get results based on something that I wrote. (Check Anna's for 'Keith P. Graham', and the first book listed is mine.)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#496
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

VC and startups are fundamentally about disruption. You can't make an omelette without breaking a few eggs (laws). The incumbent players are not going to sit still and let things be "disrupted". A common response is to make sure the public knows about the broken eggs. I would say youtube, Google, Spotify, Uber, doordash, etc. all have made my life much better.

[deleted]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#497

My ISP will shut off my internet if it catches me torrenting copyrighted material but if you're a massive corporation that steals TBs of data its barely a blip in the news.

Wouldn't it be amazing if all of Meta's ISPs cut them off for torrenting? One can dream...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#498
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

VC and startups are fundamentally about disruption. You can't make an omelette without breaking a few eggs (laws). The incumbent players are not going to sit still and let things be "disrupted". A common response is to make sure the public knows about the broken eggs. I would say youtube, Google, Spotify, Uber, doordash, etc. all have made my life much better.

You don't know a world without them so you actually have no idea if they have made your life compared to that world much better or much worse. How your life was at the time is irrelevant.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#499
post #353

Earlier quoted context omitted.

Same would be the case for YouTube. Google case was different in that AFAIK there wasn't any obvious legal problem with indexing, and, back then, they were actually doing everyone a favor. Hardly anyone had any issue with Google search until the time when news media screwed themselves over by going all in on ads, overdoing it, then trying to bring back the paywall, only to realize no one is actually browsing their si…

And then news media in Canada got even worse a few years ago. They demanded the government make a law so that when Google or Facebook even links to an article, they must pay the news org. Google decided to pay the link tax. Facebook decided to block all news links. From talking to people, most think that Facebook is the villain, in reality it's the news orgs in collusion with state power. https://en.wikipedia.org/wik…

Then Facebook complied with the new law which had negative outcomes for Canadians. The politicians then blamed Facebook for the negative outcomes: https://www.wired.com/story/meta-facebook-instagram-news-blo...

I do believe that large companies should be taxed to help improve society. This law was not the right way to do it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#500

Earlier quoted context omitted.

Yes. And the problem here isn't that companies get away with doing things like this, the problem is that individuals don't. Attempting to lock information behind a nightmarish legal system is the problem. I'm pretty much at the point now where I don't buy the "copyright incentivizes creation" argument any more. Copyright, like advertising, incentivizes creation by enormous corporations, but also like advertising it i…

Sure creative people will always create but the scope of that creativity will be limited if we do away with intellectual property. Steve Spielberg would probably always have created movies, but he wouldn't have been able to make Jurassic Park, Saving Private Ryan,or Indiana Jones without capital from the studio system, and the studio system wouldn't have provided him with that capital of they couldn't extract economi…

I think a limited, short copyright can do good that the current many-year copyright does not. Imagine a 1 year copyright in the context of film. Companies would prioritize box office sales no doubt, but that’s how it used to be and it was generally positive. It’s really the extremity of modern copyright that I think causes these issues.
Post reply on HN