Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

861–870 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#861

We developers like to pretend that LLM's are akin to humans and that they've been using things like NYTimes like humans as educational material. But they are not. It's much simpler, proprietary writing is now integrated into the source code of OpenAI, it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. Claiming copy paste is a natural evolving process of millions of ye…

>It's much simpler, proprietary writing is now integrated into the source code of OpenAI

The source code of the LLM is likely a few hundred lines of text describing the shape of the neural networks involved in the model.

None of the NYTimes content will be in the source code. NYTimes doesn't publish Python source code, it publishes human language news.

LLMs are conceptually simple, mostly matrix multiplications and some non-linear operations connecting each layer, in some loops based on attention, etc. It's the staggering amount of training data and compute that makes them complex.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#862

Earlier quoted context omitted.

We should all be worried about that. If journalism is replaced with AI, truth is replaced with the AI hallucination du jour.

It’s already done. it’s not the future.

Yup. News has been rampant with speculation, hearsay, and propaganda all my life。 Content mills and astroturfers already bury the truth or relevant stories with noise.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#863

The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…

I think NYT is going to win. LLMs are arguably compressed data archives with weird algorithms. The fact that they will regularly regurgitate verbatim quotes of training data is evidence of this, as are the guardrails that try to prevent this. The second piece of evidence is this paper explained here https://www.hendrik-erz.de/post/why-gzip-just-beat-a-large-l... where instead of an LLM researchers used gzip compresse…

Unfortunately GZIP won't beat LLMs for text classification. The research you cited is just poorly done science that has been widely debunked. The original paper compared top-2 accuracy of GZIP with top-1 accuracy with BERT. The dataset also contains a lot of train/test data leakage. See this article for the rebuttal: https://kenschutte.com/gzip-knn-paper/ and this thread for a previous discussion on hackernews: https://news.ycombinator.com/item?id=36758433.

Further, the evidence presented by NYT in the lawsuit could be hard to reproduce. I tried multiple prompts on multiple versions of GPT-4 APIs but still could not get GPT-4 to reproduce NYT articles exactly. NYT might as well tried to let GPT-4 reproduce 100,000 articles and only found a few cases where GPT-4 actually recited the whole article. In that case OpenAI might as well be arguing that this is only a rare bug and avoid losing the lawsuit in a massive way.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#864

Earlier quoted context omitted.

> How would you pay for news otherwise? You could subsidise news via "public service" style stipends. Much like having a government owned "independent" news service (eg the BBC) this comes with a high risk of corruption. Don't bite the hand that feeds and all that. You could implement a much lower friction non-recurring payment system. I'd be far more tempted to drop a little money on a fixed term (5 articles, 1 day,…

> I'd be far more tempted to drop a little money on a fixed term (5 articles, 1 day, ???) setup than a subscription. I believe that’s often referred to as a newspaper, which should be available in all good newsagents on any given day.

I'll just pop out to the nearest newsagent that stocks the NYT so I can engage with a news aggregator

(But yes, that is the model)

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#865
post #847
post #194

Earlier quoted context omitted.

What will happen in this case is that large content providers will get paid directly and smaller content providers will get rolled up into a licensing bag and get small indirect payouts. For example, we might see a model where people who's books have been used will get a pay out proportionate to the sales of the book (perhaps), so if your books sells just a few thousand copies expect $20 but if you sell millions expe…

What will happen is that all will go to China and maaaybe some third world country, or run your own models from shady sources. So you will use a Chinese AI that spies on you, or you will use some shady service from a shady country (that will play cat and mouse like torrent sites).. or most likely you will run your own model when you are computer literate and no model if you are not. Actually most models are so loboto…

This is an odd discussion as llms are really bad for authoritative information distribution, they are really untrustworthy! But, if that changes and they do start to be a reliable support as a general information assistant then I see things more like Spotify vs. Napster. I would prefer that there would be a greater diversity of sources and indirect is going to require more accountability than music but somewhat like that.

No need to emigrate!

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#866
post #793

Earlier quoted context omitted.

> machines covered in lab grown tissue I think you're misapprehending — I mean an entity fully 3D printed out of tissue, no machinery (unless you're counting all biology as machinery, but I think you're not doing that). I recon bio-printing is now where home computing was in the Apple 1 era, so this is a way off, but it's foreseeable . > The process of being human is a big part of what makes us human. Mmm. How much h…

I recon bio-printing is now where home computing was in the Apple 1 era How do you recon that, Apple 1 was Turing complete. We haven't printed life, that would be a tremendous accomplishment. I think we're closer to Edison inventing a lightbulb as a step to computers being possible. Printing a conscious thing, at all, would be like the transistor. An Apple 1 analogue wouldn't be likely because of the terrible ethics…

> We haven't printed life, that would be a tremendous accomplishment.

Sure we have, and in multiple different senses.

The ones which matters here are cell culture, which is nowhere near the fanciest bar that's been surpassed in this field, and tissue culture which is somewhat harder but the reason why I recon it's at the Apple 1 level is that a small number of experimentalists are messing around with it using expensive equipment that you can technically buy at home but you need to be well trained to actually use, for example:

https://youtu.be/Z_ZGq8Tah0k?si=u6bBatjuSWcyNYJ3

And, more broadly, there's bioprinting as a research etc. field:

https://en.wikipedia.org/wiki/3D_bioprinting

And here's a TED talk from ten years ago where they demoed an early research 3D printed kidney on stage:

https://youtu.be/bX3C201O4MA?t=11m46s

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#867
post #464

Earlier quoted context omitted.

> But I disagree with the underlying assumption that you can anthropomorphize LLMs. Gradient descent and backpropagation don't take place in the brain. LLMs "learn" in the same way that Excel sheets "learn". Backprop doesn't happen in us, but I think our neurones still do gradient descent – synapses that fire together, wire together. And ultimately, at the deepest level we can analyse, our brains' atoms are doing qua…

>Backprop doesn't happen in us, but I think our neurones still do gradient descent – synapses that fire together, wire together. No! Hebbian learning is categorically NOT gradient based learning. Hebbian update rules are local and not the gradient of any function. Cortical learning is so vastly different from how artificial neural networks “learn” they cannot even begin to be meaningfully compared mathematically. Heb…

Finally, a good counterargument. I've seen enough terrible arguments to know exactly how you feel — even in specifically just AI.

I have to keep reminding myself that outside of my own speciality, ChatGPT knows more than me despite its weaknesses, so I bet ChatGPT knows more about Hebbian learning than I do.

I'll look into that more.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#868

Earlier quoted context omitted.

I can’t think of a better way to argue in favor of “LLMs are copyright laundering machines” than from the humanness angle. Humans have rights, software tools don’t. If you grant an LLM the full set of human rights, then it can consume information, regurgitate copyrighted works, and use it to generate money for itself. However, considering blatantly obvious theft as “homage” goes hand in hand with free will, agency, b…

Humans don't exactly have the greatest track record of granting other humans rights. I don't presume they'll get it any better with AI. What I expect to happen is whoever has the most influence and power will get what they want and we'll end up raising a generation with the implicit understanding of "that's just how things are," natural order, truth, reality, and all that jazz. The only thing that ever changes outcom…

I can’t argue for or against whether LLMs should have rights or not… I can only point out the hypocrisy of claiming LLMs are “like human” enough and independent enough that their operators-become-slaveowners cannot be held to account on any copyright matters, but also claiming that LLMs are not like human at all lest someone demans them to have rights and nukes the industry.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#869

Earlier quoted context omitted.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

It's not okay for a human to pirate, plagiarize, violate IP rights and laws, etc. But I disagree with the underlying assumption that you can anthropomorphize LLMs. Gradient descent and backpropagation don't take place in the brain. LLMs "learn" in the same way that Excel sheets "learn". Humans are living beings with needs and rights. A person being able to legally squat in a home doesn't mean that a drone occupying p…

> anthropomorphize LLMs (...) gradient descent (...) backpropagation (...) needs and rights

You misunderstood me. I was talking about something more fundamental.

Understanding is data compression. They are the same thing. Learning patterns, building mental models, creating abstractions, generalizing, gaining intuition/a feel for something - all the things humans engage in as part of learning and understanding the world - are all acts of lossy data compression.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#870

I've been arguing since ChatGPT came out that LLMs should fall under fair use as a "transformative work". I'm not a lawyer and this is just my non-expert opinion, but it will be interesting to see what the legal system has to say about this.

Paywalled content as well?
Post reply on HN