Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

751–760 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#751
Verbatim usage of content is copyright infringement obviously but speaking English is not. Learning from content is not copyright infringement either. I don't know if NYT has a clause for this type of usage of their content but still I don't think it would be covered by copyright the way I understand it

As long as an LLM rephrases what it learned and not regurgitate verbatim text, it should be fine but we'll see what the judge says

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#752

Earlier quoted context omitted.

I see, the narrative switched form “cat’s out of the bag” to “genie’s out of the bottle”. Regardless, no one wants to ban llms. We just want the theft to stop.

There is no theft. Hyperbole won't get you taken seriously, use correct terminology.

There is no correct terminology. My life became a lot easier when I realized that language changes meaning not just across different languages, but words take on subtly or significantly different meanings based on culture and dialect.

A lot of red-blue state misunderstandings are based on that, as are ones across US racial subgroups. Ditto for lawyer-engineer conversations.

"Theft" has pretty different meanings depending on whom you're speaking to. Legal jargon here is quite different from business, which can be quite different from popular. That's okay!

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#753

Earlier quoted context omitted.

I used to work as a computer programmer until I retired. Nearly always, my work was part of a collaborative effort, and latterly didn't include any copyright claim. My income was never impacted by unauthorized copying. Until the 80s, there was no copyright on software, and yet even then people made a living programming. Craftsmen don't claim copyright on their artifacts. Furniture designs were widely copied; but Chip…

I very much doubt the company or foundation you were working for was selling the non-copyrighted software. If it was, it probably only worked on very specific hardware that you also produced and were selling, and thus copying it was largely useless. If you were working for a university, than the university obviously doesn't make money from selling software, and thus doesn't care for copyright as much. Also, craftsmen…

> I very much doubt the company [...] was selling the non-copyrighted software

Well you'd be mistaken. Lately, it was custom software, for a particular client, and of no interest to others. Earlier, it was before software copyright was a thing, and computer manufacturers gave software away to sell the hardware.

At the very beginning, yes, it was "very specific" hardware; it was Burroughs hardware, which used Burroughs processors. But that was before microprocessors, and all hardware was "very specific".

> (plus they rely on trademark laws and design patents quite often)

Craftsmen and labourers were earning a living long before anyone had the idea of a "trademark", still less a "design patent".

> The output of lawyers is in fact sometimes copyrighted

You're right. That's why I didn't say "lawyers", I said "legal advocates". Those are people who speak on your behalf in courts of law, not scribes writing contracts. Anyway, the ancient Greeks and Romans had written laws, contracts and so on; they managed without trademarks and copyrights.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#754
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

I'm not for or against anything at this point until someone gets their balls out and clearly defines what copyright infringement means in this context. If you give a bunch of books to a kid all by the same author and then pay that kid to write a book in a similar style and then I go on to sell that book...have I somehow infringed copyright? The kids book at best is likely to be a very convincing facsimile of the orig…

There are two problems with the “kid” analogy:

a) In many closely comparable scenarios, yes, it’s copyright infringement. When Francis Ford Coppola made The Godfather film, he couldn’t just be “inspired” by Puzo’s book. If the story or characters or dialog are similar enough, he has to pay Puzo, even if the work he created was quite different and not a literal “copy”.

b) Training an LLM isn’t like giving someone a book. Among other things, it involves making a derivative copy into GPU memory. This copy is not a transitory copy in service of a fair use, nor likely a fair use in itself, nor licensed by the rights-holder.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#756
post #99

Earlier quoted context omitted.

If you study copyrighted material for four years at a university and then go on to earn money based on your education, do you owe something to the authors of your text books? I'm not sure how we should treat LLMs with respect to publicly accessible but copyrighted material, but it seems clear to me that "profiting" from copyrighted material isn't a sufficient criteria to cause me to "owe something to the owner".

Do people ever get tired of this argument that relies on anthropomorphizing these AI black boxes? A computer isn't a human, and we already have laws that have a different effect depending on if it's a computer doing it or a human. LLMs are no different, no matter how catchy hyping them up as being == Humans may be.

> we already have laws that have a different effect depending on if it's a computer doing it or a human

which laws?

we generally accept computers as agents of their owners.

for example, a law that applies to a human travel agent also applies to a computerized travel agency service.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#757

Earlier quoted context omitted.

The issue as I see it is that every bit of data that the model ingested in training has affected what the model _is_ and therefore every token of output from the model has benefited from every token of input. When you receive anything from an LLM, you are essentially receiving a customized digest of all the training data. The second issue is that it takes an enormous amount of training data to train a model. In order…

What's wrong with paying copyright holders, then? If OpenAI's models are so much more valuable than the sum of the individual inputs' values, why can't the company profit off that margin? >That’s like a person having to pay a little bit of money to all of their teachers and mentors and everyone they’ve learned from every time they benefit from what they learned. I could argue that public school teachers are paid by p…

> What's wrong with paying copyright holders, then?

There’s nothing wrong with it. But it would make it vastly more cumbersome to build training sets in the current environment.

If the law permits producers of content to easily add extra clauses to their content licenses that say “an LLM must pay us to train on this content”, you can bet that that practice would be near-universally adopted because everyone wants to be an owner. Almost all content would become AI-unfriendly. Almost every token of fresh training content would now potentially require negotiation, royalty contracts, legal due diligence, etc. It’s not like OpenAI gets their data from a few sources. We’re talking about millions of sources, trillions of tokens, from all over the internet — forums, blogs, random sites, repositories, outlets. If OpenAI were suddenly forced to do a business deal with every source of training data, I think that would frankly kill the whole thing, not just slow it down.

It would be like ordering Google to do a business deal with the webmaster of every site they index. Different business, but the scale of the dilemma is the same. These companies crawl the whole internet.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#758
post #731
post #602

Earlier quoted context omitted.

I don't want them to steamroll everyone equally here, but to not steamroll anyone. I think you're nissing the point, and putting cart before horse. If you ensure that corporations are treated as stringently as people are sometimes, the reverse is true. And that means your goal will presumably be obtained, as the corporate might, becomes the little guy's win. All with no unjust treatment.

Huh. I see downvotes. I am mystified, for if people and corporations are both treated stringently under the law, corporations will fight to have overly restrictive laws knocked down. I envision pitting corporate body against corporate body, when one corporatism lobbies, works to (for example) extend copyrights, others will work to weaken copyright. That doesn't happen as vigilantly currently, because there is no corp…

Corporations follow these laws much more stringently than individuals. Individuals often use pirated software to make things, I've seen many examples of that. I've never seen a corporation use pirated software to make things, they pay for licenses. Maybe there is some rare cases, but pirating is mostly a thing individuals do not corporations.

So in general it is already as you say, corporations are much more targeted by these laws than individuals are. These laws mostly hinders corporations, us individuals are too small to be noticed by the system in most cases.

I've also seen indie games use copyrighted material with no issues, but AAA titles seem to avoid that like the plague. I can't really think of many examples where corporations are breaking these laws more than small individuals do.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#759

Does anyone know what the copyright status of LLM generated content is? That is, if I feed a NYT article into GPT4 and say, summarize this article, and then publish that summary, is there argument or precedent that says that is or is not copyright infringement? Asking for a friend.

There is no difference between an LLM summarizing a copyrighted work and a Wikipedia contributor summarizing a copyrighted work.

Wikipedia has some words on how summaries related to copyright law: https://en.wikipedia.org/wiki/Wikipedia:Plot-only_descriptio...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#760
post #733
post #708

Earlier quoted context omitted.

“Feeling slighted” is a gross understatement of how a lack of compensation flowing to creators has shaped the internet and the wider world over the past 25 years. If we have a problem with the way top media companies compensate their creators, that is a separate issue - not a justification for layering another issue on top.

YouTube had made way more content creators wealthy than the NYT. Writers are not going to be paid more after this ruling either way.

Has the NYT made even a single content creator wealthy? Journalists there make less money than an average software engineer.
Post reply on HN