Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

201–210 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#201
Is the answer for LLMs/OpenAI to properly cite/give credit to the authoritative source? If they did that, would NYT still have a claim/case? I’d still think yes because the content is not publicly available/behind a paywall, so some sort of different subscription/redistribution of content license would likely be appropriate ? But then after that license/agreement (which I assume they must already have something like this in place, no?) if they cite/give credit to the source instead of a rewording/summarization/claiming as it’s own, seemingly that might be enough to thwart legal challenges?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#202

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

This argument is moot. Just because some countries - see china - steal intellectual property it doesnt mean we should. There are rules to the games we play specifically so we dont end up like them.

Countless Americans are happily 'stealing' intellectual property everyday from other Americans by accessing two websites — SciHub and LibGen — who owe their very existence to them being hosted in foreign countries with weak intellectual property protection and not being subject to US long-arm jurisdiction. Even on this website, using sites like archive.is (which would be illegal if they operated in the US) to bypass paywalls to access copyrighted material is common and rarely frowned upon. I doubt a culture of respecting copyright is as characteristic of "us" as you seem to think.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#203
post #109

Interesting. I think the appropriation, privatization, and monetization of "all human output" by a single (corporate) entity is at least shameless, probably wrong, and maybe outright disgraceful. But I think OpenAI (or another similar entity) will succeed via the Sackler defense - OpenAI has too many victims for litigation to be feasible for the courts, so the courts must preemptively decide not to bother with compen…

What concerns me, and I don’t see mentioned as much as I would expect, is: how will people be compensated for generating new content if ChatGPT takes over?

I believe the innovation that will really “win” generative AI in the long term is one that figures out how to keep the model populated with fresh, relevant, quality information in a sustainable way.

I think generative AI represents a chance to fundamentally rethink the value chain around information and research. But for all their focus on “non-profit” and “good for humanity”, they don’t seem very interested in that.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#204

Earlier quoted context omitted.

To have a positive impact on the world? Also, presumably NYT still has a business model unrelated to whatever OpenAI is doing with their data and everyone working there is still getting paid for their work...

Oh thank goodness we can rely on charity for our information economy > Also, presumably NYT still has a business model unrelated to whatever OpenAI is doing with [NYT’s] data… That’s exactly the question. They are claiming it is destroying their business, which is pretty much self-evident given all the people in here defending the convenience of OpenAI’s product: they’re getting the fruits of NYTimes’ labor without p…

> Oh thank goodness we can rely on charity for our information economy

You seem to be assuming an "information economy" should exist at all. Can you justify that?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#205
post #123

Google can look up into their index and can remove whatever they want to, within minutes. But how that can be possible for an LLM? That is, "decontaminate" the model from certain parts of the corups? I can only think of excluding the data set from the training and then retrain? As a side note, I think LLM frenzy would be dead in few years, 10 years time frame at max. The rent seeking on these LLMs as of today would n…

> almost no workload (other than CAD, Graphics) runs on Windows or Unix including this very forum About a fifth to a quarter of public-facing Web servers are Windows Server. Most famously, Stack Overflow[1]. [1]: https://meta.stackexchange.com/a/10370/1424704

> About a fifth to a quarter of public-facing Web servers are Windows Server

Got a link for that? Best I can find is 5% of all websites: https://www.netcraft.com/blog/may-2023-web-server-survey/

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#206

Earlier quoted context omitted.

If NYT wins this, then there is going to be a massive push for payouts from basically everyone ever…I don’t see that wallet being fat for long.

If LLMs actually create added value and don't just burn VC money then they should be able to pay a fair price for the work of people they're relying upon. If your business is profitable only when you get your raw materials for free it's not a very good business.

What is a fair price? The entire NYT library would be a fraction of a fraction of the training set (presumably).

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#207
post #123

Google can look up into their index and can remove whatever they want to, within minutes. But how that can be possible for an LLM? That is, "decontaminate" the model from certain parts of the corups? I can only think of excluding the data set from the training and then retrain? As a side note, I think LLM frenzy would be dead in few years, 10 years time frame at max. The rent seeking on these LLMs as of today would n…

But, wasn’t the reason proprietary unixes died out at major work horses because of a nearly feature comparable free alternative (Linux)? Extending the analogy, LLMs won’t die out, just proprietary ones. (Which is where I think this tech will actually go anyway.)

LLMs won't die out but proprietary LLMs behind APIs might not have valuations of hundreds of billions of dollars.

Crowd source, crowed trained (distributed training) fast enough, good enough generative models that are updated (and downloadable) every few months would start to erode the subscriber base gradually.

I might be very very wrong here but it seems like so from where I see it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#208

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

So Chinese LLMs are bad actors, but USA LLMs are the good guys? I don't see it that way, but I'm sure from an American perspective that how it seems.

You've missed the point he was making -- that Chinese and Russian companies don't care about American copyright and will do whatever is in their interest.

And although you were being flippant, yes, Chinese LLMs are bad actors.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#209
These lawsuits could end up being a nightmare for AI companies if the plaintiffs are successful. One can’t easily just remove the content from a model like you can from a website if someone sends you a takedown notice. The content is deeply embedded inside the mathematical relationships within the model. You’d basically have to retrain the whole model again sans the offending data. Given the cost to retrain just a few successful claims would destroy any business built around making money off these models.

The times appears to have a strong case here with their complaint showing long verbatim passages being produced by ChatGPT that go far beyond any reasonable claim of fair use. This will be an interesting case to watch that could shape the whole Generative AI space.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#210
post #36

I think the train has left the station and the ship has sailed. I'm not sure it's possible to put this genie back in the bottle. I had stuff stolen by OpenAI too, and I felt bad about it (and even send them a nasty legal letter when it could output my creative work almost verbatim), but I think at this point, the legal landscape needs to somehow adjust. The Copyright Clause in the US Constitution is clear: To promote…

>the ship has sailed Certainly, but debating the spirit behind copyright or even "how to regulate AI" (a vast topic, to put it mildly) is only one possible route these lawsuits could take. I suspect that ultimately the winner is going to be business first (of course in the name of innovation), and the law second, and ethics coming last -- if Google can scan 129 million books [1] and store them without even a slap on…

The court decided to focus on the tiny snippets Google displayed rather than the full text on their servers backing the search functionality. The court found significant that Google deliberately limited the snippet view so it couldn't be used as a replacement for purchasing the original book. The opinion is a relatively easy read, I highly recommend it if you're interested in the issue. It's also notable the court commented that the Google case was right on the edge of fair use.

https://law.justia.com/cases/federal/appellate-courts/ca2/13...

Post reply on HN