Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

141–150 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#142
post #36

I think the train has left the station and the ship has sailed. I'm not sure it's possible to put this genie back in the bottle. I had stuff stolen by OpenAI too, and I felt bad about it (and even send them a nasty legal letter when it could output my creative work almost verbatim), but I think at this point, the legal landscape needs to somehow adjust. The Copyright Clause in the US Constitution is clear: To promote…

>the ship has sailed

Certainly, but debating the spirit behind copyright or even "how to regulate AI" (a vast topic, to put it mildly) is only one possible route these lawsuits could take.

I suspect that ultimately the winner is going to be business first (of course in the name of innovation), and the law second, and ethics coming last -- if Google can scan 129 million books [1] and store them without even a slap on the wrist [2], OpenAI and anyone of that size can most surely continue to do what they're doing. This lawsuit and others like it are just the drama of 'due process'.

[1] https://booksearch.blogspot.com/2010/08/books-of-world-stand... [2] https://www.reuters.com/article/idUSBRE9AD0TT/

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#143
Here's the most important part (from NYT story on the lawsuit [1]):

In one example of how A.I. systems use The Times’s material, the suit showed that Browse With Bing, a Microsoft search feature powered by ChatGPT, reproduced almost verbatim results from Wirecutter, The Times’s product review site. The text results from Bing, however, did not link to the Wirecutter article, and they stripped away the referral links in the text that Wirecutter uses to generate commissions from sales based on its recommendations.

[1] https://www.nytimes.com/2023/12/27/business/media/new-york-t...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#144
Will be interesting to see where this ends up.

If I scrape the NYT content, and then commercialize a service that lets users query that content through an API (occasionally returning verbatim extracts) without any agreement from or payment to the NYT, that would be illegal.

It's not obvious to me why putting an LLM in the middle of the process changes that.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#145

This is just rent seeking from dying media instead of working on creating something new in my view. AI indeed is reading and using material sa a source, but is deriving results based on that material. I think this should be allowed, but now it is a fight who has better paid politicians pretty much. I am open to hear other thoughts.

Here’s another thought: It’s good that there are real incentives to produce original content. Especially investigative journalism which is an extremely tough business financially — even without LLMs — but with lots of social value.

It would be silly to totally destroy the incentive to produce new technologies like LLMs, but so wouldn’t it be silly to destroy the incentive to produce original, high-quality content either for human or LLM consumption.

FWIW the LLMs are obviously the ones rent-seeking here, if you’re trying to use the term for its actual meaning instead of just “charge a subscription for something I don’t want to pay for.”

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#146

You do copyright for content that you invented and which didn't exist before. But NYT content is reporting on events truthfully to the public without any fiction or lies. Since there can be only one truth it should not matter whether NYT or Washington Post or ChatGPT is spinning it out. Unless NYT is claiming they don't report truth and publishes fiction. That is of concern since, NYT claims to reporth news truthfull…

AFAIK facts like happenings in the world are not copyrightable. So I guess the nyt is arguing it's copying their prose and way of writing about them?

Yes. Journalism is a job. People do the work of turning these happenings into words, and are paid for it. That's what's stolen here. The value created through doing that work.

If it didn't have value, Microsoft would lose nothing by no longer ingesting it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#147
If I were the CIA/US gov officials I would somehow want the NY Times to drop this case as one would not want AIs that don't have the talking points and propaganda pushed via papers not be part of the record.

I am not saying that the NY Times is a CIA asset but from the crap they have printed in the past like the whole WMDs in Iraq saga and the puff piece of Elizabeth Holmes they are far from a completely independent and propaganda free paper. Henry Kissinger would call the paper and have his talking point printed the next day regarding Vietnam. [1]

There is a huge conflict of access to government officials and independents of papers.

[1] https://youtu.be/kn8Ocz24V-0?si=kWyWXztWGjS_AJVl

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#148
When the US entered WWI, they couldn't build a plane despite inventing them. They had to buy planes from the French. Why? The Wright Brothers patent war [1]. This led to Congress creating a patent pool for avionics that exists to this day.

Honestly, I get this feeling about these lawsuits about using content to train LLMs.

Think of it this way: in growing up and learning to read and getting an education you read any number of books, articles, Web pages, magazines, etc. You viewed any number of artworks, buildings, cars, vehicles, furniture, etc, many of which might have design patents. We have such silliness as it being illegal to distribute photos commercially of the Eiffel Tower at night [2].

What's the differnce between training a model on text and images and educating a person with text and images, really? If I read too many NYT articles, am I going to get sued for using too much "training data"?

Currently we need copious quantities of training data for LLMs. I believe this is because we're in the early days of this tech. I mean no person has read millions of articles or books. At some point models will get better with substantially smaller training sets. And then, how many articles is too many as far as these suits go?

[1]: https://en.wikipedia.org/wiki/Wright_brothers_patent_war

[2]: https://www.travelandleisure.com/photography/illegal-to-take...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#149

Earlier quoted context omitted.

And I imagine that Gmail makes google very very special in this regard

FB likewise.

Exactly. FB happily gives away their ML tech like Llama because what they really care about is the data that can be used to train/tune models. The ML bits are just a commodity and not really worth much (something a new wave of ML startups have yet to realize).

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#150

Earlier quoted context omitted.

I don't think they're looking to prevent the inevitable, but rather see a target with a fat wallet from which a lot of money can be extracted. I'm not saying this in a negative way, but much of the "this is outrageous!" reaction to AI hasn't been about the building of models, but rather the realization that a few players are arguably getting very rich on those models so other people want their piece of the action.

If NYT wins this, then there is going to be a massive push for payouts from basically everyone ever…I don’t see that wallet being fat for long.

The data will have to become more curated. Exclusivity deals will probably become a thing too. Good data will be worth the money and hassle; garbage (or meh) data won't.
Post reply on HN