Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

131–140 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#131

You do copyright for content that you invented and which didn't exist before. But NYT content is reporting on events truthfully to the public without any fiction or lies. Since there can be only one truth it should not matter whether NYT or Washington Post or ChatGPT is spinning it out. Unless NYT is claiming they don't report truth and publishes fiction. That is of concern since, NYT claims to reporth news truthfull…

Not sure where you're coming from in this. A NYT article, once written is copyrighted. Using the content without attribution is at best plagiarism, and spitting it out the way the LLMs do is definitely a violation of if that copyright. Unless you're telling me ChatGPT has eyes and sources just like the NYT and is worrying events as it sees them too?

I don't understand. So if New York times reported on a new laws of physics and put as an article will became copyrighted? Nobody would be able to talk about it and has to discover it by themselves?

How is reporting on an event different from reporting on discovering a scientific law?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#132
post #93

Earlier quoted context omitted.

So Chinese LLMs are bad actors, but USA LLMs are the good guys? I don't see it that way, but I'm sure from an American perspective that how it seems.

What? This is about whether one country wants to cede a massive economic advantage to another country.

So the US should stop enforcing copyright or child labor laws because some other countries may not, giving them an economic advantage?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#133

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

This argument is moot. Just because some countries - see china - steal intellectual property it doesnt mean we should. There are rules to the games we play specifically so we dont end up like them.

The word ‘moot’ does not mean what you think it means.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#134

Earlier quoted context omitted.

> You do copyright for content that you invented and which didn't exist before. I dont think that's accurate. The Copyright Act, § 103, allows copyright protection for "compilations (of facts)", as long as there is some "creative" or "original" act involved in developing the compilation, such as in the selection (deciding which facts to include or exclude) and arrangement (how facts are displayed and in what order).

Okay. But ChatGPT doesn't spin out the fact in the same order right? So how does this stand in court?

as far as I understodd ChatGPT reproduced a word-for-word part of an NYT article(?), but not sure, didnt read the full post yet.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#135

Earlier quoted context omitted.

Sarah Silverman is claiming the same thing about her book. But I've tried really hard to get ChatGPT to output sentences verbatim from her book and just can't get it to. In fact, I can't even get it to answer simple questions about facts that are in her book but nowhere else -- it just says it doesn't know. Similarly I haven't been able to reproduce any text in the NYT verbatim unless it's part of a common quote or p…

The complaint has specific examples they got from ChatGPT. There is a precedent: There were some exploit prompts that could be used to get ChatGPT to emit random training set data. It would emit repeated words or gibberish that then spontaneously converged on to snippets of training data. OpenAI quickly worked to patch those and, presumably, invested energy into preventing it from emitting verbatim training data. It…

Hm... Why would people not just paste in sections of the book to the "raw" model in the playground (gpt instead of chatgpt) and just see if it completes the text correctly? Is the concern that chatgpt may have used the book for training data but not the original llm?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#136

Earlier quoted context omitted.

I don't think they're looking to prevent the inevitable, but rather see a target with a fat wallet from which a lot of money can be extracted. I'm not saying this in a negative way, but much of the "this is outrageous!" reaction to AI hasn't been about the building of models, but rather the realization that a few players are arguably getting very rich on those models so other people want their piece of the action.

If NYT wins this, then there is going to be a massive push for payouts from basically everyone ever…I don’t see that wallet being fat for long.

If they are determined to have broken the law then they should absolute be made to pay damages to aggrieved parties (now, determining if they did and who those parties are is an entirely unknown can of worms)

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#137

Earlier quoted context omitted.

which laws are broken exactly? it's not remotely settled law that "training an NN = copyright infringement"

The legal argument, which I'm sure you are very well aware of, is that training a model on data, reorganizing, and then presenting that data as your own is copyright infringement.

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#138
post #109

Interesting. I think the appropriation, privatization, and monetization of "all human output" by a single (corporate) entity is at least shameless, probably wrong, and maybe outright disgraceful. But I think OpenAI (or another similar entity) will succeed via the Sackler defense - OpenAI has too many victims for litigation to be feasible for the courts, so the courts must preemptively decide not to bother with compen…

What do you mean when you say "appropriation and privatization" of "all human output"?

The output is still there for anyone else to train on if they want.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#139

Earlier quoted context omitted.

Sarah Silverman is claiming the same thing about her book. But I've tried really hard to get ChatGPT to output sentences verbatim from her book and just can't get it to. In fact, I can't even get it to answer simple questions about facts that are in her book but nowhere else -- it just says it doesn't know. Similarly I haven't been able to reproduce any text in the NYT verbatim unless it's part of a common quote or p…

it is in the legal complaint - they have ten examples of direct content. I think they got very skilled people to work on producing the evidence.

Ah thank you. The examples start on page 30.

I wish they included the prompts they used, not just the output.

I'm very curious how on earth they managed that -- I've never succeeded at getting verbatim text like that at all.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#140
post #90

Earlier quoted context omitted.

An LLM in Russia can commit the same crime in Russia, and get sued in Russia. No idea about China, but I know Russia has a working legal system.

For some definitions of “working”.

Working enough that people and companies there exist, live, and are to some degree successful, yes. I've visited multiple times in the past few years and I found it to be pretty normal
Post reply on HN