Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

861–870 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#861

Earlier quoted context omitted.

a) If you don't have a robots.txt, you're indexed by default. It's opt-out, not opt-in. If you do nothing, you're being indexed.

It's an opt-out of an opt-in. If you run a webserver hosting your files, you already opted-in to people accessing that data. If you then don't go ahead an configure it properly, that's not exactly "opt-out" anymore. By default your files are not accessible to the network, you have to first opt-in to serving them.

Google makes a copy of your data and serves that data to users before they visit your site.

Also google cache allows users to get a copy of your site without visiting your site.

Why can they republish your data while we cannot? Why do we have to opt-out?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#863
post #226

Earlier quoted context omitted.

Even without copyright there are trade secrets, not to mention trademarks and patents. Maybe we could get rid of the latter, but I think we’d need to be pretty heavily into socialist utopia before considering nixing the former two!

Trademarks and patents are very different from copyright. Trademarks especially so because they aren't designed to "own" knowledge, just to prevent confusion about who made a product or what it is. "Intellectual property" is an abomination of a term because it conflates 3 separate mechanisms with differing goals, pretending that they're related in any meaningful sense. Patents protect a process. Trademarks protect id…

My country believes in this "intellectual property" thing so much that our copyright, patent and trademark act's name translates to "An Act on Intellectual and Artistic Works".

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#864
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

RIP Aaron Swartz

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#865
post #854

Earlier quoted context omitted.

There are other countries than the US though and if rightsholders wish to sue, lawsuits can happen there too. Several EU countries, Switzerland, South Korea, Japan, etc. are viable countries to sue from. Even in Japan which has a law specifically permitting training on copyrighted material you must still obtain it legally-- i.e. you must license it.

That's irrelevant. Switzerland (for example) isn't going to arrest Zuckerberg and put him in jail for this either. Nobody will. But if you're operating a site called Pirate Bay or something like that and it's not earning billions of dollars, expect countries to chase you across the globe trying to arrest you.

If there were a criminal prosecution for willful copyright infringement in some non-US country an extradition request for the relevant people is not an impossibility though, and there would be no legal reason to deny it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#866
post #796

Earlier quoted context omitted.

There's such a mass of possible works that it hardly constrains someone that if you could cast a magic spell preventing someone from distributing or accessing your particular work and then burned it, your spell would have essentially no effect-- no one would notice it and no one would be harmed. As long as discussion of a work that has published is not impeded, the public is not harmed even by these 50-years after li…

> But it's not a theft of goods, it's theft of service. What service? If somebody washes your windshield without you asking, it isn't a theft of service to not pay them. A theft of service arises from entering into an agreement and then failing to pay as stipulated in that agreement. Copyright isn't an agreement you can choose whether to participate in. Copyright is a legal enforcement system that imposes legal liabi…

But the thing here distinguishing it from the windshield thing is that there are so many possible texts that you choosing their particular text is to choose the work they've done.

You think of choosing somebody's particular text as the way of contracting him. Just as it isn't a restriction of your freedom of speech that going into restaurant and ordering a meal creates a contract to pay, so it isn't a restriction of your freedom of speech when you choose to seek out and repeat somebody's very particular text.

Why Harry Potter when you have any of hundreds of million of stories of similar sort that you could easily write yourself? When you choose that one, you choose it because it's already been prepared by somebody else, just as you choose restaurant because they've done work and have food ready for you. By choosing the one that's already written you accept that the author has done work for you.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#867
post #329

Earlier quoted context omitted.

I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.

They have no morals, therefore I shouldn't either! That'll teach 'em!

If buying is not owning, piracy is not stealing.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#868
post #796

Earlier quoted context omitted.

> But it's not a theft of goods, it's theft of service. What service? If somebody washes your windshield without you asking, it isn't a theft of service to not pay them. A theft of service arises from entering into an agreement and then failing to pay as stipulated in that agreement. Copyright isn't an agreement you can choose whether to participate in. Copyright is a legal enforcement system that imposes legal liabi…

But the thing here distinguishing it from the windshield thing is that there are so many possible texts that you choosing their particular text is to choose the work they've done. You think of choosing somebody's particular text as the way of contracting him. Just as it isn't a restriction of your freedom of speech that going into restaurant and ordering a meal creates a contract to pay, so it isn't a restriction of…

> so it isn't a restriction of your freedom of speech when you choose to seek out and repeat somebody's very particular text.

I hadn't made that claim, but I will in now that you've brought it up. Art operates as part of a discussion, the reference to and re-use of prior art is a key part of the how that happens. There are sooo many cases of copyright being used to limit the freedom of expression, that this really isn't disputable. Copyright clearly restricts speech.

> By choosing the one that's already written you accept that the author has done work for you.

No I don't, at least not in a sense that's different from the shoulders of all the people that author learned from and so on. Cultural works exist and take on roles in our cultural semiology, our memes our language without our choice. You can coose to not engage with a work, but you can't choose which works will be culturally relevant or not.

When you publish something, it becomes part of our shared culture and no-one has an inalienable right to own that. The limited rights we granted to encourage commercial creativity have already snowballed out of control and now people are blythly buying into another dramatic expansion of them.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#869

Earlier quoted context omitted.

What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre. The pendulum has swung back the other way. The only thing they have going for them now is the app based convenience, which is eroding as more "yellow cab" type traditional taxis band together and get set up with their own sort of city-specific app.

It's not really that surprising when the cities are passing laws to try to turn Uber back into the taxi cartel by e.g. making it harder for them to use part-time and on-demand contract drivers. The way you get the price down is by reducing friction, increasing flexibility and supply and taking advantage of efficiencies like people willing to do a dozen rides a week during surge pricing without making it a full-time j…

You think that workers having certainty of their work hours is a bad law? What are your work hours?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#870

Earlier quoted context omitted.

Yes. And the problem here isn't that companies get away with doing things like this, the problem is that individuals don't. Attempting to lock information behind a nightmarish legal system is the problem. I'm pretty much at the point now where I don't buy the "copyright incentivizes creation" argument any more. Copyright, like advertising, incentivizes creation by enormous corporations, but also like advertising it i…

Nothing stops you from downloading Ann’s archive and training a model on it, right? The likelihood that you, as an individual, get sued over is is virtually zero. This is what Meta tried to do, quietly download and use the data, to do research and advance their LLMs, without trying to establish any legal precedents or pick up fights.

In Germany people are sued for illegal movie downloading all the time. It's hard to imagine the companies behind that operation aren't aware that you can also download books.
Post reply on HN