Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

351–360 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#351
post #141

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

It's quite different, is it not? I don't get the analogy. These models are scanning, storing, and ingesting more material than any one human could. Not only is the method completely different, the end goal and applications are as well. The analogy basically isn't one, at all. I'm pretty upset at companies using our personal data to make gobs of money off of. I'm also upset that they're now using our knowledge work to…

Thank you for saying this. I can’t believe the amount of people calling for the end of copyright or saying that this is “holding back progress”. Why do we want big corporations to suck in all our work, put us out of work and not pay us anything? There’s nothing artificial about AIs like chatgpt, it’s all regurgitation of human knowledge.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#352

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

The NYT argument is going to be that they put up a site, own the copyright for their content and make that content available for either a human to read it for themselves, or software to index for something commonly understood as a search engine. Those terms do not entitle the training of LLMs for commercial use. Therefore, cease and desist. Oh and destroy anything that was created by violating the terms of our licens…

The LinkedIn case already proves that you cannot impose conditions on works you freely serve to the public. The data is there to anyone who sends a request (you don’t even need to be logged in) and if they do something you don’t like with it then oh well.

So if that’s the argument it’s already been argued by LinkedIn and lost.

This is one of those things where copyright holders have gotten absurdly full of themselves though. Like what you’ve said is that copyright holders have the right to impose a contract of adhesion on data that they are broadcasting into the public without any idea with whom they are even forming a contract, and that’s a facially absurd and incredibly noxious idea if you follow it to the conclusions it implies.

Copyright is about securing to the public works of significance and encouraging their creation and the way it’s become a lifetime-plus-75-year guarantee of intellectual ownership of ideas is fundamentally noxious and goes against the intent and spirit of the idea. And if that’s where the copyright regime is headed then I’d rather see chatGPT kill off copyright entirely.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#353

Earlier quoted context omitted.

IANAL is the dumbest abbreviation the internet has come up with. I believe I first observed these things on the Groklaw discussion threads discussing the SCO legal battle against the world. Not-A-Lawyer NAL instead of the full IANAL. I just had to say it.

IANAL is just "wacky" and "sexual" reddit tier humour, nothing more. It's boring any annoying.

IANAL long predates the use of the term “Reddit-tier”, probably goes back to usenet in the 80s.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#354
post #75
post #49

I don't think it's an exaggeration to say that LLMs might lead to the end of the open web, or at least a drastically reduced version of it. So much of these model's utility is in directly competing with the producers of the training data. Content creators and aggregators are seeing more and more reason to restrict and limit access, to avoid having AI companies consume all of their data and then be the ones making mon…

Let's assume that happens. How do I hedge against it? Is there a convenient way to mirror the bits of the web that are open now? Perhaps a mirror of archive/WayBack machine that could be viewed locally similar to Wikipedia dumps?

Yes, commoncrawl indexes

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#355
post #333

Earlier quoted context omitted.

I normally consider these discussion to be more about the people reading the comments than the people writing them. You've clearly made up your mind, but others presumably haven't so I think it's good he makes these arguments, even if it looks like tilting at windmills to you.

That's a good point, but I guess I was implicitly getting at that continually permutating poor analogies and hypotheticals isn't an interesting discussion. The fact of the matter here is that parties, such as OpenAI, are benefiting from others' knowledge work, protected or not, in a completely one-sided way and all for free. And I don't feel sorry for companies that need to build Trojan horse products, such as OpenAI…

Why should you get to profit in the workplace from the mental models you built in college from a commercially sourced textbook? Doesn’t Pearson actually own the knowledge you’re using on the job at $BIGCORP? Why do you get to profit off diffusion based on their work?

Someone else made the analogy - if you read a NYT article and then go do a stock trade based on what you read, didn’t you do that based on value generated by them, and why wouldn’t they own that too?

Like if you don’t want to talk analogies then talk principles, and humans are diffusion machines. When you write a term paper from sources you are simply diffusing those words into a new arrangement, but it’s still fundamentally someone else’s work. Why do you get to profit off the model that results from someone else’s work?

Copyright has mutated into this bizarre chimaera where people (like NYT) are essentially claiming ownership of ideas (and derivative works fall into a similar space) and that’s inherently in conflict with a system that is supposed to promote the creation of works. But it has turned into this bizarre shibboleth that if you came up with an idea it’s yours for life+75 years, completely yours and nobody else can work off it or remix it without crediting you. And that’s an unusual state, humanity hasn’t existed like this forever, the Berne convention is only 50 years old and already falling apart from unintended consequences.

Anyway there’s no proof that copyright benefits the small guy more than corporations. Disney squashing someone for writing a Star Wars fan fiction happens a lot more than Disney ripping off someone’s fanfic character for their series. Like patents there’s this mythos of it benefiting the small guy and that’s absolutely not how it works in the real world.

Also, NYT is a particularly egregious plaintiff here because they’re essentially just factual reporting of occurrences, which (like a phone book) are not really copyrightable in itself. You can copy a phone book without infringing copyright and you can train an AI model on a phonebook, and training an AI model on NYT in particular is basically doing that but for historical facts and occurrences. The fact that this costs money for NYT to generate is irrelevant, this is the “sweat of the brow” doctrine and was already swept aside by the phone book case. Just because you spent time/money making it doesn’t mean it’s copyrightable. A large amount of NYT comment is factual observation and tabulation and probably is not copyrightable in the first place. But separating that out is of course going to be challenging for NYT’s lawyers!

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#356

Earlier quoted context omitted.

Not just humans in general, the NYT specifically relies on the fact news cannot be copyrighted. Then claims it's articles are sacred...

IIRC the distinction is that facts can't be copyrighted, a particular arrangement of facts can, particularly if something more subjective (analysis or opinion) is included. So I can write my own article about the sky being shown to appear blue much of the time, but I can't copy someone else's article about the same subject.

How would one even go about a phonebook-style mechanical listing of facts and occurrences? You’re listing an impossibility and then saying that garden-variety connective sentences somehow make it not a factual listing.

Like yes if you copy a NYT article verbatim it’s like copying a phone book ads and all, and that’s infringement. But that’s not what a LLM does, NYT doesn’t like their content being used and summarized at all, even in a rearranged form that merely relies on the factual information included in the article. That’s what they want to get paid for, and unfortunately that’s not copyrightable and OpenAI is correct they don’t have to pay for that. NYT disagrees but again, they are kinda attempting to claim copyright on the factual information because they wrote some connective sentences between.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#357
post #208

Earlier quoted context omitted.

That doesn’t matter though. If I licensed material why should it count if I compete or not. If I watch Lebron James play and use it to develop an athletic training program, does it matter if I play baseball? Or WNBA? Or does it only matter if I play against him in the championship?

Licenses are bound to the purposes specified by the license. So go read the fine print.

Copyright doesn’t allow you to impose a contract of adhesion to viewers of material you’re broadcasting out into the public. I don’t agree to your contract just because it’s coming into my radio, you can’t create a term of service that imposes additional restrictions on what I can do with it. Copyright may apply but that doesn’t give you the right to enforce additional contracts of adhesion simply based on consuming the content.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#358
post #67

Earlier quoted context omitted.

Full text New York Times articles are available through subscription services other than nytimes.com. As an example, my local library offers full-text NYT articles through both nytimes.com and ProQuest. Notably, ProQuest's terms only explicitly ban scraping metadata and developing software or services that "compete or interfere" with ProQuest products: https://about.proquest.com/en/about/terms-and-conditions

and your local library and ProQuest are also bound by the same laws, even if they have an existing licensing agreement. from the ToS you just linked (note use of the terms "licensor" and "third party", which would be the NYT in this case): > Restrictions. Except as expressly permitted above, Customer and its Authorized Users shall not: > Remove any copyright and other proprietary notices placed upon the Service or an…

Grandparent is right that people keep conflating copyright and licensing. That term doesn’t allow you to violate the copyright of a licensor, but, it also doesn’t rule out (eg) fair use.

Fair use would have to be blocked via a license, and it’s going to be difficult to argue that someone agrees to a license merely by turning on their radio. Responding to unauthenticated internet requests with content is the internet equivalent of broadcast and similarly the LinkedIn case held that this did not allow LinkedIn to impose terms of service in a contract of adhesion in this fashion.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#359
post #117

Earlier quoted context omitted.

I think it's the most natural way it would happen. Pit powerful interests (publishers) against powerful interests (Microsoft/ClosedAI). The only surprising thing is it's taking publishers so long to notice this fresh abuse of copyright at grand scale will cost them.

Which is why OpenAIs GTM strategy involves the biggest players in each industry. That said, the outcome is unlikely - we have trained AI for more than a decade as ‘fair use’ at this point, it’s the application of the technology that is shifting the perspective, nor the act of training. Every computer vision system in the world is trained on mostly public data for example. Furthermore, the LLMs purpose is not to gener…

> we have trained AI for more than a decade as ‘fair use’ at this point

Fair use is about use. Spellchecking ML, search engine ML, etc. all different than ML that produces content.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#360

Earlier quoted context omitted.

I think there is a major qualitative difference between generative AI and search engines. Search engines index the web and point you at other people's work, along the way showing perhaps too much of that content (thus "stealing" users from the target webpage). But they don't reshuffle existing content into something apparently new and original. The "malicious" case for generative AI is that it sucks in copyrighted wo…

Indexing the copyrighted works is rehashing it. You're rehashing it into a different format that is more easily searchable by a computer. But that work is still based on other people's copyrighted works.

It doesn't produce a song if you ask for one. It points you at an existing one, with attribution and a (perhaps excessive) excerpt.

LLMs do.

Post reply on HN