Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

191–200 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#191
post #90

> if a federal judge finds that OpenAI illegally copied The Times' articles to train its AI model, the court could order the company to destroy ChatGPT's dataset, forcing the company to recreate it using only work that it is authorized to use. I'd like to see it happening but it sounds unrealistic.

If I read 1000s of of NYT articles to improve my writing skills, add then write an article of my own, is that a copyright violation?

if you had a subscription, no.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#192

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

>The general way LLMs work do not preserve content in it's original form: the ideas they contain are extracted and clustered statistically - as a Is the way LLM work relevant? I can make a shitty script that has as input Microsoft proprietary code and as output something identical in purpose but the text is completely different, I would rename names with synonyms, swap some things around etc. I am not against AIs, my…

I don't think the process matters much, but your proposed script would output something that was obviously very similar to the original.

What ChatGPT produces under normal use is not more similar to the NYT source than any other article on the same topic.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#193

Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…

Or on the converse: if those industries are unviable without copyright protection, they could go away entirely. This is a plausible path to "drop copyright entirely", just like encryption was dropped as an export-controlled technology in the late 90s. (remember the 40-bit "international" SSL?) OpenAI etc. have huge amounts of money behind them, they very well have a fighting chance in court to defend their usage of s…

> if those industries are unviable without copyright protection, they could go away entirely.

These creative industries include all of software development, music, TV, movies, books, media, art, etc. You do technically solve the problem of copyright by shutting all those down, but I'm not sure it's a solution anybody will vote for.

If you can come up with a serious alternative though, which can sustain those creative industries without requiring copyright, now is probably the best moment in all of history to seize the day and make that happen. There's going to be a big shake-up regardless, it's the perfect chance for alternative models.

Bear in mind that dropping copyright entirely doesn't just hurt Disney and Sony Music though - with no copyright the GPL and all other open-source licenses are unenforceable, anybody can copy & sell anybody else's art or design without permission, Spotify doesn't have to pay musicians even $0.01 any more, etc etc etc. It's not an easy problem.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#194

Earlier quoted context omitted.

they had to make a copy of the original to get it into their system in the first place!

In order to render that page, it probably was copied dozens of times all over my RAM. Do I owe NYT money now?

Well, are you making money off of those copies?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#195
post #158

Earlier quoted context omitted.

Consider a mega consulting firm with millions of von Neumann-like analysts. Together, they've processed the same vast data that LLMs have, but individually, none could. It's not an LLM, but its purpose is like ChatGPT: assisting clients with their tasks. If LLMs concern you due to their data processing, would a firm like this do the same?

Bringing up poor fitting analogies won't change my opinion.

I normally consider these discussion to be more about the people reading the comments than the people writing them. You've clearly made up your mind, but others presumably haven't so I think it's good he makes these arguments, even if it looks like tilting at windmills to you.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#196
post #154

Earlier quoted context omitted.

Well if use the tool to reproduce copyrighted content you are violating copyright. But that’s not the primary usecase and nobody in their right mind is arguing that. The weights are not a reproduction of the content. They are capable of it but so is a photocopier a lot more and we didn’t ban those either despite them technically being a lot more useful for violation. Nah, this is expansionist doctrine and agenda for…

Photocopiers are for personal use, training an AI is not. If you photocopy 10,000 copies of copyrighted text and starting distributing it you will get sued. It would be different if I trained my own AI, for my personal use.

businesses use photocopiers, there’s Xerox shops, etc.

Again, LLMs don’t copy so it’s not a good metaphor.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#197
post #171

Earlier quoted context omitted.

Or at least ask before scrapping/reading it.

If it’s on the open internet then why should they have to do that? How is openai training on articles fundamentally different from the wayback machine storing them? They’re just getting stored in a different form.

Copyright law includes many exceptions explicitly for libraries.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#198

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

IANAL is the dumbest abbreviation the internet has come up with. I believe I first observed these things on the Groklaw discussion threads discussing the SCO legal battle against the world.

Not-A-Lawyer NAL instead of the full IANAL.

I just had to say it.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#199

Earlier quoted context omitted.

100% this. But I doubt it will happen in the US, unfortunately.

The tech industry has sufficient money and influence for lobbying to push this one through. The media industry did the DMCA adjustments to copyright reasonably fast, and tech industry is even more powerful and wealthy.

Yeah but in this case there are extremely influential forces on both sides of the issue.

I think when this happens it is normally easier to block a law than to push it through, so I expect the current laws will remain for the short/medium term.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#200
The precedent people should be paying much more attention to is sampling in music. When it first arose, it really wasn’t clear what status it had. There was at least a decade when people basically thought it was legal to use small samples of other recordings because they were small and the new use turned them into something unrecognisably different. Which was kind of logical, actually, but turned out not to be true!

The current legal requirement to get clearance for all samples only arose after a bunch of court cases in the late 80s/ early 90s, mostly involving quite obscure musicians.

There are a lot of people on here who assume that ‘logic will prevail’ in the courts on questions like use of copyrighted data in training data. History shows that this really isn’t a safe assumption. The courts have historically been extremely favorable to copyright holders. It would be foolish to underestimate the legal risk to openai et al here

Post reply on HN