Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

691–700 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#691
post #583

Earlier quoted context omitted.

I’m not an expert in AI training, but I don’t think it’s as simple as storing writing. It does seem to be possible to get the system to regurgitate training material verbatim in some cases, but my understanding is that the text is generated probabilistically. It seems like a very difficult engineering challenge to provide attribution for content generated by LLMs, while preserving the traits that make them more usefu…

Sure, it's a hard problem, but as others have pointed out frequently in this thread.. there is not only "no incentive" to solve it but a clear disincentive. If one can say where the data comes from, one might have to prove that it was used only with permission. And the reason why it's a hard problem is not related to metadata volume being greater than content volume. Clearly a book title/year published is usually sho…

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#692
post #119

Earlier quoted context omitted.

“Through Microsoft’s Bing Chat (recently rebranded as “Copilot”) and OpenAI’s ChatGPT, Defendants seek to free-ride on The Times’s massive investment in its journalism by using it to build substitutive products without permission or payment,” the lawsuit states. I can't be the only one that sees the irony of this news being "reported" and regurgitated over dozens of crappy blogs. ChatGPT [..] “can generate output tha…

All those blogs are _also_ violating copyright, so I don't see the irony? One doesn't spend a million dollars suing a defendant with pennies to their name. I'd also expect the Times style complaint to have merit because it's probably much easier for ChatGPT to imitate the NYT style than an arbitrary style.

> an arbitrary style

..off to try "Gwern style"

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#693
post #327

Earlier quoted context omitted.

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

A human can't credit the source of each element of everything they've learnt. AI's can't either, and for the same reason. The knowledge gets distorted, blended, and reinterpreted a million ways by the time it's given as output. And the metadata (metaknowledge?) would be larger than the knowledge itself. The AI learnt every single concept it knows by reading online; including the structure of grammar, rules of logic,…

At the same time, there are situations where humans are expected to provide sources for their claims. If you talk about an event in the news, it would be normal for me to ask where you heard about it. 100% accuracy in providing a source wouldn’t be expected, but if you told me you had no idea, or told me something obviously nonsense, I would probably take what you said less seriously.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#694

Earlier quoted context omitted.

> It's ultimately a kind of lossy statistical compression scheme at some level. And on this subject, it seems worthwhile to note that compression has never freed anyone from copyright/piracy considerations before. If I record a movie with a cell phone at a worse quality, that doesn't change things. If a book is copied and stored in some gzipped format where I can only read a page at a time, or only read a random page…

If you watch a bunch of movies then go on to make your own movie based on influence from these movies, you are protected even if you have mentally compressed them into your own movie. At some point, you can learn, be influenced and be inspired from copyrighted material (not copyright infringement), and at some point you are just making a poor copy of the material (definitely copyright infringement). LLMs are probably…

There's no obvious need to hold people / AI to same standards here, yet, even if compression in mental-models is exactly analogous to compression in machine-models. I guess we decided already that corporations are already "like" persons legally, but the jury is still out on AIs. Perhaps people should be allowed more leeway to make possibly-questionable derivative works, because they have lives to live, and genuine if misguided creative urges, and bills to pay, etc. Obviously it's quite difficult to try and answer the exact point at which synthesis & summary cross a line to become "original content". But it seems to me that, if anything, machines should be held to higher standard than people.

Even if LLMs can't cite their influences with current technology, that can't be a free pass to continue things this way. Of course all data brokers resist efforts along the lines of data-lineage for themselves and they want to require it from others. Besides copyright, it's common for datasets to have all kinds of other legal encumbrances like "after paying for this dataset, you can do anything you want with it, excepting JOINs with this other dataset". Lineage is expensive and difficult but not impossible. Statements like "we're not doing data-lineage and wish we didn't have to" are always more about business operations and desired profit margins than technical feasibility.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#695
post #508

Earlier quoted context omitted.

Robot.txt isn't about copyrights, its about preventing bots. Its effectively a EULA. Copyright law only goes into effect when you distribute the content you scrape. If you scraped New York times for your own LLM that you used internally and didn't distribute the results, there would be no copyright infringement.

Er... This is what all these lawsuits against LLMs are hoping to disprove

Which lawsuits are concerning LLMs used only privately by the organization that developed it?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#696
post #678

Earlier quoted context omitted.

I don’t view LLMs as a fad. It’s like drummers and drum machines. Machines and drummers co-exist really well. I think drum machines, among other things, made drummers better.

It mainly made mediocre drummers sound better to the untrained ear.

Then it comes down to preference, but the craft and discipline objectively evolved as a result. Just as your trained ear may keep your preference to more refined percussive - a subject matter expert may care more for their native, untrained materials on their topic. In either case, music progressed in spite of the trained ears, just as AI will progress all walks of life in spite of the subject matter experts.

Nonetheless, trained ears and subject matter experts can still pick their preference.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#697
post #634

Earlier quoted context omitted.

A human faces the same restriction, if it provides commercial services on the internet creating code that is a copy of copyrighted code.

This isn't true; if you hire a contractor and tell them "write from memory the copyrighted code X which you saw before", and they have such a good memory that they manage to write it verbatim, then you take that code and use it in a way that breaches copyright, you're liable, not the person you paid to copy the code for you. They're only liable if they were under NDA for that code.

> they have such a good memory that they manage to write it verbatim

No, there is no clause in copyright law that says "unless someone remembered it all and copied it from their memory instead of directly from the original source." That would just be a different mechanism of copying.

Clean-room techniques are used so that if there is incidental replication of parts of code in the course of a reimplementation of existing software, that it can be proven it was not copied from the source work.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#698

Earlier quoted context omitted.

> if NYT goes under a dozen similar outlets can replace them overnight Not when there’s no money in journalism because the generative AIs immediately steal all content. If nyt goes under no one will be willing to start a news business as everyone will see it’s a money loser.

How does AI compete with journalism? AI doesn't do investigative reporting, AI can't even observe the world or send out reporters. Which part of journalism is AI going to impact most? Opinion pieces that contain no new information? Summarizing past events?

Sadly, opinion pieces are what drives the news economy these days. Columnists/commentators, in effect, subsidize the hard news (at those venues that even bother to produce the latter). Filling their prime time hours with opinion journalism was the trick that Fox News discovered to become wildly successful.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#699

Earlier quoted context omitted.

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

It doesn't matter what is legal. It matters what is right. Society is about balancing the needs of the individual vs the collective. I have a hard time equating individual rights with the NYT and I know my general views on scraping public data and who I was rooting for in the LinkedIn case.

When we're discussing litigation, it certainly matters what is legal.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#700

Earlier quoted context omitted.

Sarah Silverman is claiming the same thing about her book. But I've tried really hard to get ChatGPT to output sentences verbatim from her book and just can't get it to. In fact, I can't even get it to answer simple questions about facts that are in her book but nowhere else -- it just says it doesn't know. Similarly I haven't been able to reproduce any text in the NYT verbatim unless it's part of a common quote or p…

The complaint has specific examples they got from ChatGPT. There is a precedent: There were some exploit prompts that could be used to get ChatGPT to emit random training set data. It would emit repeated words or gibberish that then spontaneously converged on to snippets of training data. OpenAI quickly worked to patch those and, presumably, invested energy into preventing it from emitting verbatim training data. It…

> OpenAI quickly worked to patch those

So it was a problem, but isn't anymore?

Post reply on HN