Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…
Short of a police state, how would you enforce this? This has napster -> subscription spotify energy. But the only people happy about that are Spotify and people who found it distasteful to download music illegally. There just wasn’t a consumer-friendly option for a while, so the black market was the only market. So. The enforcement mechanism is what… a scary DMCA letter? (There will definitely be a stupid DCAIA in t…
New York Times considers legal action against OpenAI as copyright tensions swirl
331–340 of 383 posts
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#332IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…
If I read a story about a flood in Dubai on the NYTimes and then I write an email to my friend summarizing what I just read, it is not copyright infringement. I am not sure why it would suddenly become infringement because an LLM is composing that email for me.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#333Earlier quoted context omitted.
Bringing up poor fitting analogies won't change my opinion.
I normally consider these discussion to be more about the people reading the comments than the people writing them. You've clearly made up your mind, but others presumably haven't so I think it's good he makes these arguments, even if it looks like tilting at windmills to you.
The fact of the matter here is that parties, such as OpenAI, are benefiting from others' knowledge work, protected or not, in a completely one-sided way and all for free. And I don't feel sorry for companies that need to build Trojan horse products, such as OpenAI and Google, in order to survive off of other people's data that they never compensate for.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#334Earlier quoted context omitted.
Violating TOS, at least to scrape and use later, is legal.[0] I'm not sure how the ruling interacts with LLMs, but I'm sure OpenAI's lawyers would bring it up. [0]: https://www.forbes.com/sites/zacharysmith/2022/04/18/scrapin...
FYI LinkedIn actually won that case after appealing once more: https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn
I see they went to the Supreme Court who kicked it back to the Ninth who then re-affirmed their position that HiQ Labs was not in violation of the CFAA.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#335Earlier quoted context omitted.
I don't know why people use these analogies. No person can memorize terabytes worth of lyrics.
Does that invalidate the analogy? The point I'm trying to make is if you rule these behaviors illegal, then you're necessarily making intelligent AI illegal, because humans are capable of the same behaviors.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#336Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…
Copyright largely remains about the PRODUCTION of content, not about the CONSUMPTION of it.
Someone who grew up reading Marvel comics being able to make new original comics in that style is perfectly ok. That same person perfectly replicating an Avengers comic is going to land them in hot water.
The focus on infringement really needs to be on what LLMs produce, not their training.
There definitely needs to be something like a secondary pass added which checks output against a vectordb of the training set to avoid too close derivative IP outputs (and ideally checks for jailbreaking or inappropriate content at the same time).
Any production services would need to subscribe to a service like that to stave off litigation on infringement, much like how the oft repeated complaints regarding YouTube copyright infringement eventually dissipated as content tagging was added (and shifted to complaints over too broad application of it).
A generative AI model having read the NYT but producing new original news articles in the style of a newspaper is a very weird argument for infringement.
A human driven service or an automated one that takes current NYT articles and summarizes or reworks them, publishing itself and cutting them out of the ad revenue is more problematic (but also widespread already and generally considered protected).
Services which exactly duplicate their articles would be more clearly infringement, but there's no evidence that's even a fraction of what ChatGPT is doing.
Criminalizing training would set back whatever county did so significantly in global competition for a critical new economic (and defense) trend, and would ultimately only be a minor stop gap for copyright holders as you'd simply see a market for secondhand generated content from foreign models trained on copyrighted data but then producing content that was itself not copyrightable but could be used to train domestic models.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#337Earlier quoted context omitted.
>>"If you do allow that, the many many affected industries have catastrophic problems." That is the problem. Technically, AI should be allowed to 'read' content, it isn't hidden, and it gets mixed with other content in a 'brain' like thing. AI and Humans can both spit out a new product that is 'similar' and thus be sued on that similarity. But it can also produce endless similar variations at low cost and fast. It is…
There’s an unbelievably vast difference between a human’s creative process and the mechanical reproduction of reweighed training data. Machines don’t create, people do.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#338Earlier quoted context omitted.
So I can use an open source LLM like Llama then?
You've pierced my completely precise, absolutely airtight choice of language about this situation as some sort of flaw in the greater point being made. Less glibly: a non-profit oriented LLM is just in a little different place on the scale, but doesn't fundamentally change my takeaway. However in this situation it makes it particularly egregious.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#339Earlier quoted context omitted.
FYI LinkedIn actually won that case after appealing once more: https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn
Where do you see that they won the case? Can you provide a source because the wikipedia article directly contradicts what you are saying...? I see they went to the Supreme Court who kicked it back to the Ninth who then re-affirmed their position that HiQ Labs was not in violation of the CFAA.
[0] https://www.natlawreview.com/article/court-finds-hiq-breache...
[1] https://www.natlawreview.com/article/hiq-and-linkedin-reach-...
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#340Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…
There's a headline out today about several ex-Google Brain engineers, including a co-author of “Attention Is All You Need”, setting up shop in Tokyo. [1] That's not a coincidence.
> Amid rising questions about the fairness and legality of using publicly available information to train AI models, Japan affirmed that machine learning engineers can use any data they find.
> What’s new: A Japanese official clarified that the country’s law lets AI developers train models on works that are protected by copyright.
> How it works: In testimony before Japan’s House of Representatives, cabinet minister Keiko Nagaoka explained that the law allows machine learning developers to use copyrighted works whether or not the trained model would be used commercially and regardless of its intended purpose. [2]
IANAL so don't know what the implications of this are when it comes to cross-border copyright enforcement. It's hard to imagine Japan rolling back this type of legal safe harbor. It's a boon for attracting AI startups from elsewhere and giving the local tech industry a competitive boost.
What's to stop other venues looking to grow their tech industry to do something similar? And if they do, would it create a race to the bottom type of dynamic at the expense of copyright holders?
[1] https://www.bloomberg.com/news/articles/2023-08-17/ex-google...
[2] https://www.deeplearning.ai/the-batch/japan-ai-data-laws-exp...