Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

91–100 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#91
post #60
post #36

Earlier quoted context omitted.

Do you think OpenAI doesn’t subscribe? I too load creative works into all sorts of temporary structures in order to read the paper. I don’t need to license it, I pay for a subscription. Should I pay more if I memorize the paper? Should I pay more if I read it to my sick friend in the hospital? Should I pay more if I save copies to my own hard drive and grep for words in the files? Should robots have a higher subscrip…

a subscription doesn't give you an automatic escape hatch out of copyright law. here's their ToS, which is pretty clear about what you cannot do: https://help.nytimes.com/hc/en-us/articles/115014893428-Term... (relevant parts below) Without NYT’s prior written consent, you shall not: ... (2) use robots, spiders, scripts, service, software or any manual or automatic device, tool, or process designed to data mine or sc…

Violating their terms of service doesn't really matter in terms of copyright law though. The statistical properties of that text (which is what the engine snarfs in) aren't protected by copyright.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#92
post #90

> if a federal judge finds that OpenAI illegally copied The Times' articles to train its AI model, the court could order the company to destroy ChatGPT's dataset, forcing the company to recreate it using only work that it is authorized to use. I'd like to see it happening but it sounds unrealistic.

I think it's the most natural way it would happen. Pit powerful interests (publishers) against powerful interests (Microsoft/ClosedAI). The only surprising thing is it's taking publishers so long to notice this fresh abuse of copyright at grand scale will cost them.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#93

Earlier quoted context omitted.

You don’t need a fair use exemption for transient copies in service of licensed or fair uses. Computers and networks have been around a long time. These issues have been given a good workout.

But the copy used during training is itself transient, so all this boils down really to the question of whether training a machine is fair use. Which can't be answered here exactly because the concept of fair use is deliberately vague, so this will boil down to a lawsuit and probably go to the Supremes. The USA will work something out that's reasonable as they always do and, lacking AI companies and often the concept…

It's not even a question of "fair use" because OpenAI isn't providing anybody a copy. Probability models really just aren't copyrightable to begin with.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#94

Earlier quoted context omitted.

Did the Times grant a license to every router on the internet to transmit its intellectual property to other routers? If not, the judge should grant an injunction contingent on requiring the Times to verify that every person who accesses their content is doing so only over routers and other devices with express written authorization, for every step in the process. Maybe even extend it to browsers and client libraries…

OpenAI is pretty clearly using their work to make derivative content that in certain cases (CNET) is a direct competitor. Honestly, this seems open and shut

It's not derivative though. For derivation you have to literally point to sequences of words in the original that are also in the alleged infringer, and those sequences have to be long or unique enough to not be able to come from somewhere else or just common English usage.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#95
post #72

Earlier quoted context omitted.

Google does the same to produce a search index.

you can easily opt out of that, or control it to your heart's desire (including what snippets to show), and it will be honored. There is no way to opt out of this bullshit. So no, not the same at all.

That's a courtesy of the search engine, not a requirement of copyright.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#96
post #44

Earlier quoted context omitted.

If I buy a license to listen to a Taylor Swift mp3, I can copy it onto a flash drive (or anywhere I like to use it). And that’s fine until I distribute copies to others.

Yes, and if OpenAI had a license to use the Times ’ works as model training material this would be a non-issue, of course.

There's really not a "license" in that sense for text. There's not some magical way you can use copyright to protect a probability model of the words you make, because that's style and style is not copyrightable (nor are the underlying facts the NYT reports).

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#97

Earlier quoted context omitted.

Personally, I think the Times has a far better case for presenting mechanical copies (including with mechanical alteration) than it does with model training. The tool building of model training is more likely to be fair use than the use of the tool to provide mechanical copies of copyirght-protected material that competes directly with the original in the market.

Fair use only covers very limited circumstances which probably does not include selling a subscription (ChatGPT+). If you’re selling a repackaged reproduction of someone else’s copyrighted works, that’s never protected by fair use.

But you can't copyright the underlying facts nor the style of writing, just the literal words used in your article. If you can't point to a sequence of words in the original that's in the later work, that later work isn't "derivative".

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#98
post #49

I don't think it's an exaggeration to say that LLMs might lead to the end of the open web, or at least a drastically reduced version of it. So much of these model's utility is in directly competing with the producers of the training data. Content creators and aggregators are seeing more and more reason to restrict and limit access, to avoid having AI companies consume all of their data and then be the ones making mon…

Is this open web in the room with us? Cheekiness aside, the myth of beautiful open web always seems to be exagerated. Most content on the web is already generated, ugly and spammy. Most of the traffic is already owned by mega corporations. One could easily argue that it's unlikely it'll get worse - if anything, AI could empower competition as now a group of 3 passionate, free writers can compete with agenda-driven, f…

> Most content on the web is already generated, ugly and spammy.

Now? Yes. I grew up when it was just ugly, and by ugly I mean animated gif backgrounds with obvious seams on the tile boundary.

*old man shakes fist at The Cloud*

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#99
post #59

Earlier quoted context omitted.

ignorance of copyright law won't save you here. i can't legally torrent a copywrited music file just because "no tos to agree to" when i listen to it.

There is no copyright involved in training a model. No precedent at least. Only online ToS/API restrictions exist for scraping content. The DMCA issues exist entirely on the (re)distribution side of coyright material. So seeding in a torrent swarm = redistribution. If you download some copyright material somehow and don't share it with anyone, there is no caselaw that says anything about it.

RIAA kept running into this problem which is why it only sent the lawyers after people they could confirm seeded.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#100
These mega LLMs that can autonomously roam the web and consume original content are basically the "I made this" meme[0] and having some legal precedent would be good for all users of the web.

[0] - https://knowyourmeme.com/memes/i-made-this

Post reply on HN