Any time I see the phrase "democratize access," my spidey-sense starts tingling. It's almost never used to describe an action that's an unadulterated good for society. It's USUALLY used to describe something sketchy at best, or even outright evil, with the justification that only "the bad guys" have access currently, and everything would be better if EVERYONE had access. Look, I get that unethical corporations using…
I've developed a similar mistrust of the term. Blogged about it a while back https://writing.kemitchell.com/2023/01/05/Type-Error-Democra...
The Battle over Books3
111–120 of 135 posts
Re: The Battle over Books3
#112(We noticed The Pile was recently taken offline, so we hosted it: https://thenose.cc. Apparently Books3 was also a part of The Pile, so feel free to download.)
Re: The Battle over Books3
#113Earlier quoted context omitted.
They were, usually to preserve the work, not distribute it on a large scale because there was no large scale. What's your argument ?
This isn’t the case. In Ancient Rome it was a common practice for someone in the audience to note down poets’ performances, then give that transcript to a team of amanuenses who would produce copies for sale, with none of the proceeds going back to the original creator. In the pre-copyright world, no one saw any problem with this practice; as the other poster mentioned, the creator economy was patronage-based. All th…
Re: The Battle over Books3
#114Earlier quoted context omitted.
That's how I see it as well. To democratize art, music etc. now means to remove the skill component with the usage of all the work done so far. No one is actually prevented from pursuing those things and if you don't want your art to become training data you're ruining democratization and are somehow against the will of the people.
The level of disrespect for artists I've seen here and on other forums with regards to this technology has been staggering. The entitlement for work you didn't do is so gross.
Re: The Battle over Books3
#115>He sees the widespread practice of training AI on copyrighted data as outrageous, and finds it infuriating that this behavior gets defended with claims that it’s democratizing access to information. “Open source doesn’t mean you took a bunch of people’s shit and gave it away for free,” he says. “That's theft.” >Whether the defendant had purchased a signed copy or flagrantly shoplifted a dog-eared paperback wouldn’t…
Writing a review or a book report is very much creating a derivative work in copyright law. Copyright law then says that these derivative works are a fair use. It does not follow that other derivative works that you personally feel are as serious are also fair use.
Re: The Battle over Books3
#116Earlier quoted context omitted.
Good luck when everything is paywalled
Given the already huge cost of training, and the evident lack of concern the LLM folks seem to have for copyright, why wouldn't the AI groups purchase subs to scrape the paywalled content? The would possibly need to apply some effort to appear human, but that should only throttle the rate, not stop their scraping all together.
Clearly, places like Reddit have wised up to this and are making API usages non-free for example, so while it's not impossible, you can see the limitations being put into place already. Twitter is another one.
It seems like all this data is now considered gold and people lock up gold?
Re: The Battle over Books3
#117Earlier quoted context omitted.
It's pretty much the opposite: Big corporates are the ones training the models (and benefiting from it).
Other big corporation hold a large portion of copyrights trained on as well. I tend to side with everyone being able to do it
Re: The Battle over Books3
#118Earlier quoted context omitted.
books3.tar.gz itself is ~37gb compressed. Often really the entire "The Pile" dataset (composed of both the mostly compressed archives, along with a ~450gb compressed jsonl compilation of the data) is being discussed. That's around 825gb.
That is shockingly approachable for a large fraction of English literature.
Re: The Battle over Books3
#119Earlier quoted context omitted.
Have you never seen the outcome of a class action? They’re all slaps in the wrist, less than speeding tickets, and the action members get like a free hotdog or red bull as compensation if they’re lucky
yes, the lawyers always clean up the aim would be to make training on copyrighted material legally toxic and render all existing datasets and trained weights unlawful the damages are simply a bonus
"We bought these weights in good faith from Digitus Tertius corp of bermuda."
Re: The Battle over Books3
#120Books3 is the easy/fast/cheap method but if the quality of the model really brings in some sort of revenue is there anything stopping a company from buying/checking out of the library all these books, scanning and ORC'ing them and adding to the model the hard way?