Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

581–590 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#581
post #423
post #385

Earlier quoted context omitted.

What makes you think AI researchers (including the big labs like OpenAI and Anthropic) aren't trying to solve these problems?

the solutions haven't arrived. neither have changes in lieu of having solutions. "trying" isn't an actual, present, functional change. and it just gets passed around as an excuse for companies to keep doing whatever they're doing.

Please recall how much the world changed in just the last year. What would be your expected timescale for the solution of this particular problem and why is it more important than instilling models with the ability to logically plan and answer correctly?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#582

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

"probably the single most important development in human history" is the kind of hyperbole you'd only find here. Better than medicine, agriculture, electrification, or music? That point of view simply does not jive with what I see so far from AI. It has had little impact beyond filling the internet with low-effort content. I feel like the crypto evangelists never got off the hype train. They just picked a new destina…

> Better than medicine, agriculture, electrification, or music?

Shoulders of giants.

Thanks to the existence of medicine, agriculture, and electrification (we can argue about music), some people are now healthy, well fed, and sufficiently supplied with enough electricity to go make LLMs.

> I hope the NYT is compensated for the theft of their IP and hopefully more lawsuits follow.

Personally I think all these "theft of IP" lawsuits are (mostly) destined to fail. Not because I'm on a particular side per-se (though I am), but because it's trying to fit a square law into a round hole.

This is going to be a job for legislature sooner or later.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#583

Earlier quoted context omitted.

A human can't credit the source of each element of everything they've learnt. AI's can't either, and for the same reason. The knowledge gets distorted, blended, and reinterpreted a million ways by the time it's given as output. And the metadata (metaknowledge?) would be larger than the knowledge itself. The AI learnt every single concept it knows by reading online; including the structure of grammar, rules of logic,…

> And the metadata (metaknowledge?) would be larger than the knowledge itself. Because URLs are usually as long as the writing they point at?

I’m not an expert in AI training, but I don’t think it’s as simple as storing writing. It does seem to be possible to get the system to regurgitate training material verbatim in some cases, but my understanding is that the text is generated probabilistically.

It seems like a very difficult engineering challenge to provide attribution for content generated by LLMs, while preserving the traits that make them more useful than a “mere” search engine.

Which is to say nothing about whether that challenge is worth taking on.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#584

I'd bet they win, but how do you possibly measure the dollar amount? If you strip out 100% of NYT content from GPT-4, I don't think you'd notice a difference. But if you go domain by domain and continue stripping training data, the model will eventually get worse.

I think the only feasible outcome of the NYT winning would be a royalty structure that would have OpenAI paying the NYT to access their work, including back payments

I am not saying that NYT will win, but I think it is more likely to win because it has many more supporters (including politicians, judges) than developers do.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#585
post #314

Earlier quoted context omitted.

Playing back large passages of verbatim content sold as your “product” without citation is almost certainly not fair use. Fair use would be saying “The New York Times said X” and then quoting a sentence with attribution. Thats not what OpenAI is being sued for. They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. This is also related…

> They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. In what sense are they claiming their generated contents as their own IP? https://www.zdnet.com/article/who-owns-the-code-if-chatgpts-... > OpenAI (the company behind ChatGPT) does not claim ownership of generated content. According to their terms of service, "OpenAI hereby assig…

They are distributing the output, so they (implicitly) claim to have the right to distribute it. I can send you a movie I downloaded along with a license that says "I hereby assign to you all our right, title, and interest, if any, in and to Output. ", I'm still obviously infringing on the copyright of that movie (unless I have a deal that allows re-distribution, of course, as Netflix does).

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#586

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

> I hope this results in Fair Use being expanded to cover AI training. Couldn't disagree more strongly, and I hope the outcome is the exact opposite. I think we've already started to see the severe negative consequences when the lion's share of the profits get sucked up by very, very few entities (e.g. we used to have tons of local papers and other entities that made money through advertising, now Google and Facebook…

So you want all the profit to be sucked up by the three companies that can afford to make deals with rights holders to slurp up all their content?

Making the process for training AI require an army of lawyers and industry connections will have the opposite effect than you intend.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#587
post #307

Earlier quoted context omitted.

Copyright law is far from perfect, but the concept is not morally bankrupt. It is certainly abused by large entities but it also, in principle, protects small content creators from exploitation as well. In addition to journalists, writers, musicians, and proprietary software vendors, this also includes things like copyleft software being used in unintended ways. When I write copyleft software, it is my intention that…

With the exception of source code availability, copyleft is mostly about using copyright to destroy itself. Without copyright (which I feel is unethical), and with additional laws to enforce open sourcing all binaries, copyleft need not exist. So it is not good when people use copyleft as a justification for copyright, given that its whole purpose was to destroy it.

Source code availability (and the ability to modify the code on a device) is the most important part, IMO , regardless of RMS's original intention. Do you feel that it's ethical that OpenAI is keeping their model closed?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#588
It seems weird to sue an AI company because their tool "can recite [copyrighted]" content verbatim.

If I paid a human to recite the whole front page of the New York Times to me, they could probably do it. There's nothing infringing about that. However, if I videotape them reciting the front page of the New York Times and start selling that video, then I'd be infringing on the copyright.

The guy that I paid to tell me about what NYT was saying didn't do anything wrong. Whether there's any copyright infringement would depend what I did with the output.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#589
post #421

Earlier quoted context omitted.

Everyone learns from papers. That's the point of them, isn't it? Except we pay, what, $4 per Sunday paper or $10/mo for the digital edition? Why should a robot have to pay much more just because it's better at absorbing information?

Because the issue isn’t the intake, it’s the output, where your analogy breaks down. If you could clone the brain of someone who was “trained” on decades of NYT and could reproduce its information on demand at scale, we’d be discussing similar issues.

Your analogy doesn't make sense either.

If we could clone the brain of someone I hardly think we'd be discussing their vast knowledge of something so insignificant as the NYT. I don't think we should care that much about an AI's vast knowledge of the NYT either or why it matters.

If all these journalism companies don't want to provide the content for free they're perfectly capable of throwing the entire website behind a login screen. Twitter was doing it at one point. In a similar vein, I have no idea why newspapers are complaining about readership while also paywalling everything in sight. How exactly do they want or expect to be paid?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#590
post #375

Earlier quoted context omitted.

There's a few levels to this... Would it be more rigorous for AI to cite its sources? Sure, but the same could be said for humans too. Wikipedia editors, scholars, and scientists all still struggle with proper citations. NYT itself has been caught plagiarizing[1]. But that doesn't really solve the underlying issue here: That our copyright laws and monetization models predate the Internet and the ease of sharing/paywa…

Can you imagine spending decades of your life, studying skin cancer, only to have some $20/month ChatGPT index your latest findings and spit out generically to some subpar researcher: "Here's how I would cure melanoma!" followed by your detailed findings. Zero mention of you. F-that. Attribution, as best they can, is the least OpenAI can do as a service to humanity. It's a nod to all content creators that they have b…

If the future AI can indeed cure disease my mission of working in drug discovery will be complete. I’d much rather help cure people (my brother died of melanoma) than protect any patent rights or copyrighted text.
Post reply on HN