Earlier quoted context omitted.
How is this relevant? >The RIAA accused her of downloading and distributing more than 1,700 music files on file-sharing site KaZaA Emphasis mine. I think most people would agree that whatever AI companies are doing with training AI models is different than sending verbatim copies to random people on the internet.
> I think most people would agree that whatever AI companies are doing with training AI models is different than sending verbatim copies to random people on the internet. I think most artist who had their works "trained by AI" without compensation would disagree with you.
US Copyright Office found AI companies breach copyright. Its boss was fired
241–250 of 410 posts
Re: US Copyright Office found AI companies breach copyright. Its boss was fired
#242Earlier quoted context omitted.
They never said model training is a violation of copyright. The ruling says model training on copyrighted material for analysis and research is NOT copyright infringement, but the commercial use of the resulting model is: "When a model is deployed for purposes such as analysis or research… the outputs are unlikely to substitute for expressive works used in training. But making commercial use of vast troves of copyrig…
The vast trove of copyright work has to refer to training. ChatGPT is likely on the order of 5-10TB in size. (Yes, Terabyte). There are college kids with bigger "copyright collections" than that...
Disk size is irrelevant. If you lossy-compress a copyrighted bitmap image to small JPEG image and then sell the JPEG image, it's still copyright infringement.
Re: US Copyright Office found AI companies breach copyright. Its boss was fired
#243Earlier quoted context omitted.
If AI is so important, maybe it should be owned by the government and free to use for all citizens.
Name two non-military things that the government owns and aren't complete dumpster fires that barely do the thing they're supposed to do. Even (especially?) the military is a dumpster fire but it's at least very good at doing what it exists to do.
Library of Congress
National Park Service
U.S. Geological Survey (USGS)
NASA
Smithsonian Institution
Centers for Disease Control and Prevention (CDC)
Social Security Administration (SSA)
Federal Aviation Administration (FAA) air traffic control
U.S. Postal Service (USPS)
Re: US Copyright Office found AI companies breach copyright. Its boss was fired
#244Earlier quoted context omitted.
How is this relevant? >The RIAA accused her of downloading and distributing more than 1,700 music files on file-sharing site KaZaA Emphasis mine. I think most people would agree that whatever AI companies are doing with training AI models is different than sending verbatim copies to random people on the internet.
> I think most people would agree that whatever AI companies are doing with training AI models is different than sending verbatim copies to random people on the internet. I think most artist who had their works "trained by AI" without compensation would disagree with you.
[1] used purely as an example
Re: US Copyright Office found AI companies breach copyright. Its boss was fired
#245One aspect that I feel is ignored by the comments here is the geo-political forces at work. If the US takes the position that LLMs can't use copyrighted work or has to compensate all copyright holders – other countries (e.g. China) will not follow suit. This will mean that US LLM companies will either fall behind or be too expensive. Which means China and other countries will probably surge ahead in AI, at least in t…
Well hell, by that logic average citizens should be able to launder corporate intellectual property because China will never follow suit in adhering to intellectual property law. I'm game if you are.
Musicians remain subject to abuse by the recording industry; they're making pennies on each dollar you spend on buying CDs^W^W streaming services. I used to say, don't buy that; go to a concert, buy beer, buy merch, support directly. Nowadays live shows are being swallowed whole through exclusivity deals (both for artists and venues). I used to say, support your favourite artist on Bandcamp, Patreon, etc. But most of these new middlemen are ready for their turn to squeeze.
And now on top of all that, these artists' work is being swallowed whole by yet another machine, disregarding what was left of their rights.
What else do you do? Go busking?
Re: US Copyright Office found AI companies breach copyright. Its boss was fired
#246Earlier quoted context omitted.
>I'm not sure I understand your question. It's reasonably clear that transformers get caught reproducing material that they have no right to. The kind of thing that would potentially result in a lawsuit if you did it by hand. Is that a problem with the tool, or the person using it? A photocopier can copy an entire book verbatim. Should that be illegal? Or is it the problem that the "training" process can produce a mo…
Let's start with I think a case that everyone agrees with. If I were to take an image, and compress it or encrypt it, and then show you data file, you would not be able to see the original copyrighted material anywhere in the data. But if you had the right computer program, you could use it to regenerate the original image flawlessly. I think most people would easily agree that distributing the encrypted file without…
Those extra steps are meaningfully different. In your description, a casual observer could compare the two JPEGs and recognize the inferior copy. However, AI has become so advanced that such detection is becoming impossible. It is clearly voodoo.
Re: US Copyright Office found AI companies breach copyright. Its boss was fired
#247Earlier quoted context omitted.
>It was snark intended to point out that corporations don't get to have their cake and eat it too. "have their cake and eat it too" allegations only work if you're talking about the same entity. The copyright maximalist corporations (ie. publishers) aren't the same as the permissive ones (ie. AI companies). Making such characterizations make as much sense as saying "citizens don't get to eat their cake and eat it too…
Yes they are. Look at what happened when deepseek came out. Altman started crying and alleging that deepseek was trained on OpenAI model outputs without an inkling of irony
Can you link to the exact comments he made? My impression was that he was upset at the fact that they broke T&C of openai, and deepseek's claim of being much cheaper to train than openai didn't factor in the fact that it requried openai's model to bootstrap the training process. Neither of them directly contradict the claim that training is copyright infringement.
Re: US Copyright Office found AI companies breach copyright. Its boss was fired
#248Earlier quoted context omitted.
How is this relevant? >The RIAA accused her of downloading and distributing more than 1,700 music files on file-sharing site KaZaA Emphasis mine. I think most people would agree that whatever AI companies are doing with training AI models is different than sending verbatim copies to random people on the internet.
Who knew alls she needed was to change the tempo, pitch, timbre, add/remove lyrics, add/subtract a few notes, rearrange harmony, put it behind a web portal with a fancy name, claim it had an inspirational muse or assume all mortal beings as being without one in the first place so it doesn't matter, and proceed to make millions off of said process methodically rather than giving it away for free, and she'd be right as…
Re: US Copyright Office found AI companies breach copyright. Its boss was fired
#249Earlier quoted context omitted.
"Rearranging text" is not what modern LLMs do though, unless you specifically ask them to.
I didn't make this claim. Feel free to bring a cogent argument to a commenter who did.
???
Did you not literally comment the following?
>A new research paper is obviously materially different from "rearranging that text to create a marginally new text".
What did you mean by that, if that's not your claim?
Re: US Copyright Office found AI companies breach copyright. Its boss was fired
#250Earlier quoted context omitted.
It's only complete non-sense if you understand how humans learn. Which we don't. What we do know though is that LLMs, similar to humans, do not directly copy information into their "storage". LLMs, like humans, are pretty lossy with their recall. Compare this to something like a search indexed database, where the recall of information given to it is perfect.
Well, you don't get to pick and choose in which situations an LLM is considered similar to a human being and in which not. If you argue that it similarly to a human is lossy, well let's go ahead and get most output checked by organizations and courts for violations of the law and licenses, just like human work is. Oh wait, I forgot, LLMs are run by companies with too much cash to successfully sue them. I guess we jus…
Another way would be to train an internal model directly on published works, use that model to generate a corpus of sanitary rewritten/reformatted data about the works still under copyright, then use the sanitized corpus to train a final model. For example, the sanitized corpus might describe the Harry Potter books in minute detail but not contain a single sentence taken from the originals. Models trained that way wouldn't be able to reproduce excerpts from Harry Potter books even if the models were distributed as open weights.