Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

671–680 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#671

Earlier quoted context omitted.

That's irrelevant. The main point is that they are re-distributing the content without permission from the copyright owners, so they are sort of implicitly claiming they have copy/distribution rights over it. Since they don't, then it's obvious they can't give you this content at all.

>The main point is that they are re-distributing the content without permission from the copyright owners, By your logic, Firefox is re-distributing content without permission from the copyright owners whenever you use it to read a pirated book. ChatGPT isn't just randomly generating copyrighted content, it just does so when explicitly prompted by a user.

That is not the same thing at all. If I search on Google for copyrighted content and Google shows me the content, it is the server which serves the content who is most directly responsible, not Google nor I. Firefox is only a neutral agent, whereas ChatGPT is the source of the copyrighted content.

Of course, if the input I give to ChatGPT is "here is a piece from an NYT aricle, please tell it to me again verbatim", followed by a copy I got from the NYT archive, and ChatGPT is returning the same text I gave it as input, that is not copyright infringement. But if I say "please show me the text of the NYT article on crime from 10th January 1993", and ChatGPT returns the exact text of that article, then they are obviously infringing on NYT's distribution rights for this content, since they are retrieving it from their own storage.

If they returned a link you could click, t and retrieved the content from the NYT, along with any other changes such as advertising, even if it were inside an iframe, it would be an entirely different matter.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#672

Earlier quoted context omitted.

> that is not free Why did you specify that this stuff you like, you only like if it's "not free"? The hidden assumption is that the information you like wouldn't be made available unless someone was paying for it. But that's not in evidence; a lot of information and content is provided to the public due to other incentives: self-promotion, marketing, or just plain interest. Would you prefer not to have access to Wik…

I’ll restate it for clarity: I like high-quality information. Producing and publishing high-quality information is not free. There are ways to make it free to the consumer , yes. One way is charity (Wikipedia) and another way is advertising. Neither is free to produce; the advertising incentive is also nuked by LLMs; and I’m not comfortable depending on charity for all of my information. It is a lot cheaper to produc…

Well, its existence does prove it's possible!

I contribute to Wikipedia, and I don't consider my contributions to be "charity"; I contribute because I enjoy it. Even in the age of printing presses, copyright law was widely ignored, well into the 20thC. The USA didn't join the Berne Convention until 1989 (and they promptly went mad with copyright).

Yes, there's only one Wikipedia; but there are lots of copies, and lots of similar efforts. Yes, there's one Wikipedia, like there's one Mona Lisa. There are lots of things of which there's only one; in that sense, Wikipedia isn't remotely unique.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#673

Earlier quoted context omitted.

Except Google forced these companies to use their platform (Google's AMP) to host the content and essentially blackmailed into doing so ("we'll link directly, but only on page 3 of results").

AMP did not need to be hosted by Google.

When it launched it absolutely did.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#674
post #641

Earlier quoted context omitted.

> the first being at the birth of modern search engines. Why do you say that? Search engines would at least direct the viewer to the source. NYT gets 35%+ of its traffic from Google: https://www.similarweb.com/website/nytimes.com/#traffic-sour...

That doesn’t mean that it wasn’t theft of their content. The internet would be a very different place if creator compensation and low friction micropayments were some of the first principles. Instead we’re left with ads as the only viable monetization model and clickbait/misinformation as a side effect.

I don't quite get it. If listing your link is considered as theft, HN is then a thief of content too. If you don't want your content stolen, just tell Google to not index your website?

I guess it's more constructive to propose alternatives than just bashing the status quo. What's your creator compensation model for a search engine? I believe whatever being proposed is trading off something significant for being more ethic.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#676

Earlier quoted context omitted.

> GPT4 is absolutely capable of providing links to sources and citations. Do you mean in the Browsing Mode or something? I don't think it is naturally capable of that, both because it is performing lossy compression, and because in many cases it simply won't know where the text that was fed to it during training came from.

[flagged]

The ability to cite some rfcs is, to me, vastly different from being able to link to sources.

By far the wildest piece of this stuff is that it near completely obliterates any traces of where the outputs come from. The black box is trained, and yes sometimes some salient data pole rfc's are captures, but generally where each training comes from is not stored. That would largely defeat the purpose, would make the data it's crunching essentially incompressible, to store so much origin information.

Deeply unimpressed by this answer. This isn't linking it's sources, of where this response was trained upon. It probably got the write up & links from hundreds of other places.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#677

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

Anthropic made $200M in 2023 and projected to make $1B in 2024. That's a laggard <2 year old startup. I don't think LLMs are a fad.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#678

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

I don’t view LLMs as a fad. It’s like drummers and drum machines. Machines and drummers co-exist really well. I think drum machines, among other things, made drummers better.

It mainly made mediocre drummers sound better to the untrained ear.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#679

Earlier quoted context omitted.

I’ll restate it for clarity: I like high-quality information. Producing and publishing high-quality information is not free. There are ways to make it free to the consumer , yes. One way is charity (Wikipedia) and another way is advertising. Neither is free to produce; the advertising incentive is also nuked by LLMs; and I’m not comfortable depending on charity for all of my information. It is a lot cheaper to produc…

Well, its existence does prove it's possible! I contribute to Wikipedia, and I don't consider my contributions to be "charity"; I contribute because I enjoy it. Even in the age of printing presses, copyright law was widely ignored, well into the 20thC. The USA didn't join the Berne Convention until 1989 (and they promptly went mad with copyright). Yes, there's only one Wikipedia; but there are lots of copies, and lot…

> I contribute to Wikipedia, and I don't consider my contributions to be "charity"; I contribute because I enjoy it.

Does your personal satisfaction pay the server bills too?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#680

Earlier quoted context omitted.

Copyright isn't what got in the way here. AI could have negotiated a license agreement with the rights holder. But they chose not to.

From their perspective they're training a giant mechanical brain. A human brain doesn't need any special license agreement to read and learn from a publicly available book or web page, why should a silicon one? They probably didn't even consider the possibility that people'd claim that merely having an LLM read copyrighted data was a copyright violation.

I was thinking about this argument too: is it a "license violation" to gift a young adult a NYT subscription to help them learn to read? Or someone learning English as second language? That seems to be a strong argument.

But it falls apart because kids aren't business units trained to maximize shareholder returns (maybe in the farming age they were). OpenAI isn't open, it's making revolutionary tools that are absolutely going to be monetized by the highest bidder. A quick way to test this is NYT offers to drop their case if "open" AI "open"-ly releases all its code and training data, they're just learning right? what's the harm?

Post reply on HN