Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

881–890 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#882
post #583

Earlier quoted context omitted.

> But I prefer looser intellectual property rights anyway so Im ok with it I think more people, potentially anyways, would feel similar to to this if it applied even somewhat equally. Instead, companies can seemingly do whatever they please whereas lawyers will send letters to your home for downloading a single episode of game of thrones.

> Instead, companies can seemingly do whatever they please whereas lawyers will send letters to your home for downloading a single episode of game of thrones. I don't get it. All these companies took copyrighted data when they were tiny grew to be large, they still do that now. Google and OpenAI don't send these letters. They're not the copyright holders. I have no idea what argument you're trying to make. Corporatio…

His argument is that it's effectively a legal moat now that protects monopoly. Like we shouldn't accept that you need to break the law to have a chance to compete with them.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#883
post #808
post #782

Earlier quoted context omitted.

So... because he releases parts of his work under Creative Commons means that it is okay for him to have copyright over his books? I don't get the logic.

You can buy his work DRM-free, and much of it is public: https://www.gutenberg.org/ebooks/author/3826 https://archive.org/details/cory-doctorow-content I happily pay him for hardcovers though. I do not believe copyright makes any sense in the global post-internet world. It is a major hindrance to progress. The many countries that do not enforce copyright law will share everything that can be shared and are going to p…

You're entitled to your own opinion, but please make a difference between "free", "no-DRM" and "copyright".

And tell me how Doctorow sees it when you buy one of his book and start selling copies of it under your name.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#884
post #812
post #744

Earlier quoted context omitted.

Which is completely off topic. You can buy a paper book and own it, but it doesn't mean that you are allowed to make copies of it and sell them.

If you cannot copy it, alter it, and share it as you see fit, you do not own it.

Well when you buy a book, you own the copy of the book. Not the author's life work.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#885
post #868

Earlier quoted context omitted.

But the thing here distinguishing it from the windshield thing is that there are so many possible texts that you choosing their particular text is to choose the work they've done. You think of choosing somebody's particular text as the way of contracting him. Just as it isn't a restriction of your freedom of speech that going into restaurant and ordering a meal creates a contract to pay, so it isn't a restriction of…

> so it isn't a restriction of your freedom of speech when you choose to seek out and repeat somebody's very particular text. I hadn't made that claim, but I will in now that you've brought it up. Art operates as part of a discussion, the reference to and re-use of prior art is a key part of the how that happens. There are sooo many cases of copyright being used to limit the freedom of expression, that this really is…

But we're talking about extremely direct copying. Actual computerized copying, typically verbatim.

Doing things relating to discussion of a work are typically permitted, but you have no reason to use anybody's particular work other than to make use of the work he did in creating it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#886
post #883
post #808

Earlier quoted context omitted.

You can buy his work DRM-free, and much of it is public: https://www.gutenberg.org/ebooks/author/3826 https://archive.org/details/cory-doctorow-content I happily pay him for hardcovers though. I do not believe copyright makes any sense in the global post-internet world. It is a major hindrance to progress. The many countries that do not enforce copyright law will share everything that can be shared and are going to p…

You're entitled to your own opinion, but please make a difference between "free", "no-DRM" and "copyright". And tell me how Doctorow sees it when you buy one of his book and start selling copies of it under your name.

Impersonation would be a dick move, and would ruin my reputation. Do not need laws to avoid such things. Obvious social repercussions are enough.

Still impersonation sucks, and it happens. Thankfully we can solve this with cryptography without trying to beg the legal system of hundreds of countries to agree on enforcement tactics.

I publish 100% of code I write as FOSS. I also sign my commits. If the code shows up later without attribution to me, I would prove it publicly to call out dishonest behavior.

I would also never use legal action for this though. All information should be free. I only put FOSS licenses on code to ensure I do not get sued and so corporations bound by such silly rules have a difficult time using my work in private codebases without paying me for an alternate license.

I would abolish all IP law if I could. Let all information be free without legal risk to authors or those that share it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#887
post #335

Earlier quoted context omitted.

What are we actually worried about happening? Are AI-written books getting published? If they start out-competing humans, is that bad? According to most naysayers, they can't do anything original. Are people asking the AI for books? And then hoping it will spit it out a human-written book word for word?

I think the concern goes to the point of copyright to begin with, which is to incentive people to create things. Will the inclusion of copyrighted works in llm training (further) erode that incentive? Maybe, and I think that's a shame if so. But I also don't really think it's the primary threat to the incentive structure in publishing.

i wrote a book and copyright was not once on my mind. having created something is the incentive to create for most artists

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#888
post #425

Earlier quoted context omitted.

This really has fuck-all to do with copyright though, correct? If you can't tell how the content is before you read it, it could be written by a monkey.

This is starting to get pretty circular. The AI was trained on copyrighted data, so we can make a hypothesis that it would not exist - or would exist in a diminished state - without the copyright infringement. Now, the AI is being used to flood AI bookstores with cheaply produced books, many of which are bad, but are still competing against human authors.

shop at a real bookstore, they don't have this problem.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#889
post #884
post #812

Earlier quoted context omitted.

If you cannot copy it, alter it, and share it as you see fit, you do not own it.

Well when you buy a book, you own the copy of the book. Not the author's life work.

And I disagree with this. IP law should not exist. If I buy a thing it should be mine to do with as I see fit.

The more people that share and copy data the better.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#890
post #867

Earlier quoted context omitted.

They have no morals, therefore I shouldn't either! That'll teach 'em!

If buying is not owning, piracy is not stealing.

Whereas i agree that the current regime of "licensing" is not good, I simultaneously find it incredibly selfish to believe that one has the right to any content one likes.

"Own nothing" is bad, but so is "access and share anything." Both positions are too extreme.

Post reply on HN