Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

161–170 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#162

Earlier quoted context omitted.

If you're using their content to generate more content, doesn't it fall under fair use?

In the case of Google, merely indexing content is not considered fair use. It's a double edged sword, as media outlets have realized, after all Google is responsible for a larger part of the success of these publications. But in Germany for example, Google News basicially just copypasted articles into their service, and monetizing it without involving the publishers. That doesn't qualify as transformative even under…

The trend with Google and other search engines over the past ten years has been for them to incorporate more and more content on their own pages. It's hard to remember that not so long ago Google search results pages were just lists of web pages bereft of any other content.

Today, if you Google for a song lyric, that lyric appears in Google. You get a tiny grey source link to Musixmatch or whatever but why would anyone bother to go there if you have the complete lyric right there on the page you are looking at?

More and more content real estate has appeared on search engines' pages. Answer boxes answering questions, again with a source link few people will use, and of course the large Knowledge Graph panel filled with Wikipedia content (written by volunteers and monetised by the world's richest tech companies).

The result is that it's tech companies and their platforms that make most of the money off content (YouTube is another example). They are the oil companies of today, and like the latter use all sorts of lobbying to make sure things are organised to their advantage.

For all the ingenuity and usability they offer, they behave at least in part like parasites. They should be forced to spread the wealth round a bit more.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#163

Earlier quoted context omitted.

If I read 1000s of of NYT articles to improve my writing skills, add then write an article of my own, is that a copyright violation?

The problem with the "a lossy mathematical translation of its inputs is exactly like a person learns" arguments, even if courts don't find them ludicrous, is that people absolutely can and are found guilty of trademark violations when they read thousands of pages of the LOTR and then write a fantasy novel full of Tolkein's character names for profit.

A fantasy novel with Tolkien's characters' names is an evident copyright violation regardless of how it was generated. That's not what's happening here.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#164

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Yes, and if we reproduce something significant that way we have (depending on licences etc.) committed copyright infringement. The definition of significant is a very grey hairy one though – there are many obvious examples where you would say it is definitely copying from memory (so plagiarism rather than research or accidental similarity) or copying directly, and many examples where we'd all agree that the result is so likely to be independently produced that it isn't an issue, but even more examples where there is room for interpretation and disagreement.

This isn't as well-defined for humans as you might think, so can't be well-defined for LLM techniques by comparing them to human agents.

My argument against the current hoovering up of data under various licences for AI training, which they claim can#'t reproduce anything verbatim, is CoPilot. If there is no risk, then why did they only use public repositories and none of their own private ones? Surely they think their code contains good training material, unless they think their own code is gobbledygook. Or back in terms of licences: if it can't breach the GPL family, then it can't breach their own commercial licensing arrangements.

Would OpenAI have a problem with humans for using some of their code/documents/other in this way?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#165

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

>The general way LLMs work do not preserve content in it's original form: the ideas they contain are extracted and clustered statistically - as a

Is the way LLM work relevant? I can make a shitty script that has as input Microsoft proprietary code and as output something identical in purpose but the text is completely different, I would rename names with synonyms, swap some things around etc.

I am not against AIs, my opinion is that if your AI uses GPL code the output should be GPL, if it uses public domain images the output should be public domain images.

I mean for code if AI is actual intelligent you should be able to train an AI with C with just a few books and not with the entire GitHub open source code (and notice MS did not trained copilot on the proprietary code they have access proving they are not confident that they are in the right).

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#166

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Well, there is also the fact that if you train LLMs on LLM output the quality degrades very quickly. It is not a good thing to do.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#167

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

I tend to agree with you, but, one could argue “statistical collection of words” is a form of compression? For example, you can’t write a kids version of a novel and sell that without dealing with copyright.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#168

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Not just humans in general, the NYT specifically relies on the fact news cannot be copyrighted. Then claims it's articles are sacred...

IIRC the distinction is that facts can't be copyrighted, a particular arrangement of facts can, particularly if something more subjective (analysis or opinion) is included.

So I can write my own article about the sky being shown to appear blue much of the time, but I can't copy someone else's article about the same subject.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#169
post #146

Earlier quoted context omitted.

If I read 1000s of of NYT articles to improve my writing skills, add then write an article of my own, is that a copyright violation?

A human is not a machine.

Philosophical Mechanism and other philosophical views with Cartesianist roots would beg to differ.
Post reply on HN