Earlier quoted context omitted.
Not just crypto bros, this kind of thing is rife in politics too. Brexit is full of it. People pick one article about one minor thing in one niche area of the economy and use it to 'prove' their entire agenda.
Not just politics. I remember this being a realisation as a teenager, noticing that if you bring five reasons, the person you're talking to will refute a random one in a funny way and now the audience will decide you were wrong. Danny was the person who was absolutely the best at this. I should have written one of them down, as I can't even reproduce it but he'd use some logical fallacy to make his case which is, for…
Japan’s government will not enforce copyrights on data used in AI training
321–330 of 426 posts
Re: Japan’s government will not enforce copyrights on data used in AI training
#322Art students study art to learn how to create art, and that's completely fine, but AI models are not allowed to study art to learn how to create art, because it's copyright infringement.
Madness.
Re: Japan’s government will not enforce copyrights on data used in AI training
#323Earlier quoted context omitted.
> The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. Lossy or not, the training data provides value . If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist. If you are going to use someone else's work in order to make something that you are going to…
> If you are going to use someone else's work in order to make something that you are going to profit off of, I believe that original author should be compensated. And should also be able to decide they don't want their work used in that way. > Note that I'm not talking about what existing copyright law says; I'm talking about how I believe we should be regulating this new facet of the industry. Is it really new? Hum…
"Humans" being the important word here. I don't understand why people keep trying to compare training a model to humans learning through reading etc. They are very different things. Learning done by machines at enormous scale and done to benefit private companies financially is not the same as humans learning.
Re: Japan’s government will not enforce copyrights on data used in AI training
#324Earlier quoted context omitted.
AI models will make 1:1 copies of training data where artists try and avoid doing so. It’s common to obscure this copying by intentionally inserting lossy steps, but making an MP3 isn’t a new work. It’s most obvious when large blocks of text are recreated, but the core mechanism doesn’t go away simply because you obscure the underlying output. “Extracting Training Data from Large Language Models” https://arxiv.org/ab…
Inserting lossy steps seems to work pretty well though. https://twitter.com/giannis_daras/status/1663710057400524800...
Re: Japan’s government will not enforce copyrights on data used in AI training
#325I don't understand the AI training copyright debate. Art students study art to learn how to create art, and that's completely fine, but AI models are not allowed to study art to learn how to create art, because it's copyright infringement. Madness.
An artists studying and copying/integrating other people's art in their own style can get a job for $75,000/year.
An LLM copying/integrating everyone's data and reselling can become the most profitable company in human history, and capture the most value of every incremental piece of human generated content, in perpetuity.
Re: Japan’s government will not enforce copyrights on data used in AI training
#326I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…
The collection of copyright works for the explicit purpose of processing them for a for-profit ML model has not been shown to be fair use, and the fact that many are being marketed as for profit products that meaningfully compete with the original works is a strike against them being fair use.
Re: Japan’s government will not enforce copyrights on data used in AI training
#327I don't understand the AI training copyright debate. Art students study art to learn how to create art, and that's completely fine, but AI models are not allowed to study art to learn how to create art, because it's copyright infringement. Madness.
The difference is scale. An artists studying and copying/integrating other people's art in their own style can get a job for $75,000/year. An LLM copying/integrating everyone's data and reselling can become the most profitable company in human history, and capture the most value of every incremental piece of human generated content, in perpetuity.
Re: Japan’s government will not enforce copyrights on data used in AI training
#328Earlier quoted context omitted.
Not just crypto bros, this kind of thing is rife in politics too. Brexit is full of it. People pick one article about one minor thing in one niche area of the economy and use it to 'prove' their entire agenda.
Not just politics. I remember this being a realisation as a teenager, noticing that if you bring five reasons, the person you're talking to will refute a random one in a funny way and now the audience will decide you were wrong. Danny was the person who was absolutely the best at this. I should have written one of them down, as I can't even reproduce it but he'd use some logical fallacy to make his case which is, for…
That’s because the audience is usually at the stage of a 16 year old, even if they are older. The aim of politics and pr are easily manipulated folks.
You know, those referred to as the “the market” or “the electorate”.
Re: Japan’s government will not enforce copyrights on data used in AI training
#329Earlier quoted context omitted.
Not just politics. I remember this being a realisation as a teenager, noticing that if you bring five reasons, the person you're talking to will refute a random one in a funny way and now the audience will decide you were wrong. Danny was the person who was absolutely the best at this. I should have written one of them down, as I can't even reproduce it but he'd use some logical fallacy to make his case which is, for…
whose Danny?
Re: Japan’s government will not enforce copyrights on data used in AI training
#330I now routinely introduce this technology as "copyright laundering" and the hype put out by start-up boards and VCs as a ploy to disguise this fact. The "AI threat" is smoke-and-mirrors to dress up what's happening. I derive a huge amount of value from chatgpt because I can copy/paste without any IP impact. I could always have done this: from github, from ebooks, from many sources. Now I can benefits from the labour…