Earlier quoted context omitted.
I find this such a strange remark on this front. You got less than 1% of a book... from an author who has passed away... who wrote on a research topic that was funded by an institution that takes in hundreds of millions of dollars in federal grants each year... I'm not an author (although I do generate almost exclusively IP for a living) and I think this is about as weak a form of this argument as you possibly make.…
It isn't that someone was hurt. We have one private entity gaining power by centralizing knowledge (which they never contributed to) and making people pay for regurgitating the distilled knowledge, for profit. Few entities can do that (I can't). Most people are forced to work for companies that sell their work to the higher bidder (which are the very entities mentioned above), or ask them to use AI (under the conditi…
Meta torrented & seeded 81.7 TB dataset containing copyrighted data
761–770 of 981 posts
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#762Earlier quoted context omitted.
His criminality is one matter, but the full weight of the Federal Government on him was an entirely separate matter. A federal prosecutor's job is to jail you regardless of whether it is for downloading a file from a server or for trafficking in humans, and they will come at you with the same vigor regardless of the crime. And nothing has changed about that.
Treating a human trafficker and someone who downloaded some files from a server the same is not in the job description of a prosecutor. What an absurd statement. It's very much the job of a prosecutor to make judgements about the severity of the crime and how to respond. And in this case, the prosecutor showed incredibly poor judgement. There wasn't even a particular reason why the case should go federal in the first…
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#763Earlier quoted context omitted.
What are we actually worried about happening? Are AI-written books getting published? If they start out-competing humans, is that bad? According to most naysayers, they can't do anything original. Are people asking the AI for books? And then hoping it will spit it out a human-written book word for word?
I think the concern goes to the point of copyright to begin with, which is to incentive people to create things. Will the inclusion of copyrighted works in llm training (further) erode that incentive? Maybe, and I think that's a shame if so. But I also don't really think it's the primary threat to the incentive structure in publishing.
Copyright was invented by publishers (the printing guild) to ensure that the capitalists who own the printing presses could profit from artificial monopolies. It decreases the works produced, on purpose, in order to subsidize publishing.
If society decides we no longer want to subsidize publishers with artificial monopolies, we should start with legalizing human creativity. Instead we're letting computers break the law with mediocre output while continuing to keep humans from doing the same thing.
LLMs are serving as intellectual property laundering machines, funneling all the value of human creativity to a couple of capitalists. This infringement of intellectual property is just the more pure manifestation of copyright, keeping any of us from benefitting from our labor.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#764Earlier quoted context omitted.
Do you think developing countries are just peppered with libraries, and their inhabitants order books from Amazon? Libgen originated in Russia, and its users are global. This is not a purely American issue.
I was responding to a comment arguing that LibGen is the largest collection of knowledge in human history, which I think is an overly romantic and totally incorrect take. It may be a very useful collection of knowledge to people in developing countries, but it simply is not larger than the collection of knowledge accessible via any first-world public library. Obviously not everyone has access to that, but again, that…
How do you know or can quantify this? At a first approximation, libraries are finite in space while the internet is (for the purposes of this discussion) infinite. I'd agree with you if you had said something like Wikipedia was bigger than libgen (and probably not even then, as Wikipedia is merely a summary of primary sources, which would be theoretically contained in libgen).
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#765Earlier quoted context omitted.
I think it’s fine to criticize the hypocrisy of viciously defending the copyrights you own, while gleefully running roughshod over the ones you don’t. But it’s also possible that copyright as a concept, or in its current implementation, is bad and unjust. I’m sure some copyright holders would like nothing more than to see an argument that elevates copyright violation to the level of murder, morally or legally. But I…
the reform needs to happen at the layer where whether a copyright is valid or not is decided upon, not before (at the point of "should copyright exist") and not after (enforcement). a world without copyright means those with the largest advertising budgets will reap nearly all the rewards from new IP created by small artists. BigCorp Inc. can just sit around and wait for talented musicians to post something interesti…
This makes it sound like the majority of people produce more content than they consume.
The reality is that 99.99999% of people do not produce "art", let alone with the intention of living of it.
Whatever harms you might envision for the tiny minority who do want to try living off copyright, those concerns are dwarfed by the benefits for the rest of us.
Further, not many people who are serious about reform are literally "advocating against all copyright" A reform that simply curbed the duration to something less insane than 150 years would resolve much of what makes copyright bad, even if it continued to exist.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#766For some misterious reason I can't see Zuckerberg in front of a judge facing 50 years imprisonment. Anyone can? I truly hope that whoever takes the case goes after Meta with 1000 times the pressure that was put on Swartz, but honestly I don't expect much just as the top comment precisly expressed. And if we are going to be fair please also let's not forget about the other usual suspects, or anyone thinks they are fal…
Several EU countries, Switzerland, South Korea, Japan, etc. are viable countries to sue from. Even in Japan which has a law specifically permitting training on copyrighted material you must still obtain it legally-- i.e. you must license it.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#767Meta does a lot of stuff I disagree with, but they're usually not just straight breaking the law.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#768Earlier quoted context omitted.
“There is no ethical consumption under capitalism”
Who even thinks that ethical consumption exists under any system? Any of your consumption denies it to others. Some consumption is a necessity of course. We wouldn't speak of something absurd like "ethical breathing".
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#769We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.
I think most of the public is probably in favor of stronger IP laws now that big corps are threatening to make them jobless with IP-disrespecting AIs
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#770Earlier quoted context omitted.
It's a stupid situation, though. There are many creators I'm happy to support - but for 99% of them, I don't want their stupid merch . It's mostly low-quality garbage with high markup, that nevertheless cost something to design and produce, thus wasting both precious resources and labor - an useless tax on contributions to artists that doesn't even help anything. I really wish this wasn't necessary. (Even the okay-qu…
What is the point of this comment? Just a stream of consciousness for a future LLM sweep? Nobody thinks that the actual creator should get nothing. Are you asking for better T shirts? Do you want more direct ways of just giving cash?