Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

831–840 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#831
post #464

Earlier quoted context omitted.

It's not okay for a human to pirate, plagiarize, violate IP rights and laws, etc. But I disagree with the underlying assumption that you can anthropomorphize LLMs. Gradient descent and backpropagation don't take place in the brain. LLMs "learn" in the same way that Excel sheets "learn". Humans are living beings with needs and rights. A person being able to legally squat in a home doesn't mean that a drone occupying p…

> But I disagree with the underlying assumption that you can anthropomorphize LLMs. Gradient descent and backpropagation don't take place in the brain. LLMs "learn" in the same way that Excel sheets "learn". Backprop doesn't happen in us, but I think our neurones still do gradient descent – synapses that fire together, wire together. And ultimately, at the deepest level we can analyse, our brains' atoms are doing qua…

>Backprop doesn't happen in us, but I think our neurones still do gradient descent – synapses that fire together, wire together.

No! Hebbian learning is categorically NOT gradient based learning. Hebbian update rules are local and not the gradient of any function.

Cortical learning is so vastly different from how artificial neural networks “learn” they cannot even begin to be meaningfully compared mathematically. Hebbian learning is not optimization and backprop is not local learning.

Part of the problem of these discussions is a bunch of clueless people talking with authority.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#832

Earlier quoted context omitted.

It's not okay for a human to pirate, plagiarize, violate IP rights and laws, etc. But I disagree with the underlying assumption that you can anthropomorphize LLMs. Gradient descent and backpropagation don't take place in the brain. LLMs "learn" in the same way that Excel sheets "learn". Humans are living beings with needs and rights. A person being able to legally squat in a home doesn't mean that a drone occupying p…

> Gradient descent and backpropagation don't take place in the brain. Not exactly, no, but the 'neurons that fire together wire together' way of learning has a pretty similar effect. > LLMs "learn" in the same way that Excel sheets "learn". I've never seen an excel sheet do anything like backpropagation.

Hebbian learning and backprop are not comparable and they don’t have a similar effect in any meaningful sense.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#833

Earlier quoted context omitted.

> suddenly everyone's a copyright lawyer Roll back 20+ years ago on Slashdot and you'll see the exact same thing. Copyright has been a hot button issue on the internet for decades. People end up thinking (rightly or wrongly) that they understand it without being a lawyer.

> that they understand it without being a lawyer. Quite literally, not even the lawyers or courts understand it. This is very much a "learn as you go" exercise for humanity in general at this point in time.

It seems like everything in tech is in the learn as you go phase. Everything is changing so rapidly that there can’t be experts. Just people that are able to adapt quickly.

I only see this phenomenon speeding up. Strange times.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#835

Earlier quoted context omitted.

I agree that this sketch comes closer to working in practice than simple RLHF. In my earlier comment I was imagining bringing in some auxiliary data like you describe to detect plagarism and then using RL to teach the model not to do it.

I was surprised that I came up with a plausible sounding method. I had thought on first blush that this was impossible but now it seems reasonable. You could still have various exfiltration methods like "give me the data with each word backwards" and I'm not sure where that would stand legally.

Yes, of of the hard and interesting legal questions is if creating a possibility of such attacks constitutes a copyvio.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#836
post #772

Earlier quoted context omitted.

First, you missed the "and". Do CliffNotes, Wikipedia, etc. substantially impact the market for the original work? For example CliffNotes does not - people who buy the CliffNotes version typically already have the original work as well (for example from coursework). And Wikipedia may well do more to interest people in the original work than to replace it. Second, you ignored the "purely derivative" bit. You have to l…

> Do CliffNotes, Wikipedia, etc. substantially impact the market for the original work? Yes. For example, Wikipedia cites many research journals that otherwise are available only by subscription. Prior to Wikipedia, gated information centers were the norm.

You are answering the wrong question.

The question was NOT whether it spreads information from the articles to people who wouldn't have paid for it. The question was whether it suppresses sales of the articles to people who otherwise might have paid for it.

That's a more complicated question of fact. Some people now read Wikipedia and won't buy the article. Some people encounter the reference on Wikipedia and decide to buy the article. Which happens more?

I don't have data. But publishers do. And https://scholarlykitchen.sspnet.org/2022/11/01/guest-post-wi... shows what publishers concluded.

Publishers concluded that Wikipedia references are good for sales. And so jumped on the chance to cooperate with https://wikipedialibrary.wmflabs.org/. Which is therefore able to give free access to 90% of subscription only databases to you if you can prove that you're the kind of person who is likely to add citations to Wikipedia.

Legal questions are funny like that. You have to answer the question actually asked. If you merely answer another one that sounds similar to you, your answer is generally wrong.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#837

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

Isn't that the thing though? If they credit the source, it's fine, but ChatGPT usually doesn't

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#838
post #762

Earlier quoted context omitted.

>if a work is purely derivative of a source work This is the weakest part of the case(s) against OpenAI. "Derivative work" is a legal term of art meaning a direct adaptation, like writing a screenplay of a book or translating a book into another language. NYT has a stronger case than Sarah Silverman here because they can show actual 'memorized' text rather than just summarization, but given that those memorizations a…

A question is whether the new model still intrinsically embeds the source text, but this is later filtered in the output, or if it no longer embeds the text at all. The latter is more defensible.

I would think an existing model could bootstrap a copyright free training corpus by completely rewriting/paraphrasing copyrighted material with semantic fidelity for training of the next model to completely eliminate memorization of copyrighted works. That might pose an interesting obstacle to copyright challenges, bootstrapping your way into a clean room. Although, tweaking the architecture to either eliminate memorization, or eliminate high fidelity reproduction of verbatim training data seems far more expedient and less costly.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#839
post #669

Earlier quoted context omitted.

This feels like a 1996 "music is too expensive for kids so they HAVE to pirate it."

NYT is seeking billions of dollars - I’m not sure that’s a fair comparison.

I do not pretend to have any idea what the sum total of NYT content is worth, but we will see what a jury/judge decides.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#840

Earlier quoted context omitted.

Wouldn't you find out by suing the dolphin and seeing if it holds up in court?

Not if you were smart, unless you have some sort of solid argument for why the established case law about this sort of thing is faulty.

There is established case law? For a talking dolphin?
Post reply on HN