Earlier quoted context omitted.
As previously said, search engines index and provide links. I’ll add that it constitutes fair use because a search engine isn’t itself a replacement for the articles that it indexes. But ChatGPT is actually providing an alternative that obviates the original articles themselves.
Search engines provide links, but also titles and snippets of the page -- enough for you to decide if you want to visit, and Google will show you their cached page if you ask for it. Even the link is a copyrightable item -- artistic effort went into creating it
New York Times considers legal action against OpenAI as copyright tensions swirl
361–370 of 383 posts
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#362Earlier quoted context omitted.
Ruling in favor of copyright will call into question search engines and the like as well. Do you think Bing or Google are going to negotiate copying rights with the world's websites? LLMs are proving that intellectual property has a bunch of holes in it. It's been unstable ground to defend since day one. Upon what principle should we believe that one can own an idea and all performances or derivatives of it? Patents…
Ruling in favor of copyright will call into question search engines and the like as well. No, they won't. Search engines have already fought and won this battle on fair use grounds because they make use of the copyrighted content differently than LLMs do. It's an absolutely fundamental distinction. Patents and trademarks haven't really helped as much as they were expected to. Patents have been a thing for nearly a mi…
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#363Earlier quoted context omitted.
>which the press is now framing as “trying to hide the use of copyrighted data” Yea, now I can't read the paper and talk about it to other people it seems. The Right to Read was a prophecy I guess?
An individual or group of individuals doing this and sharing their views/summary vs. a profit-oriented program funded by major technology companies scraping this information and spitting it back out algorithmically does seem different to me. Yes, perhaps both things are on the same "sliding scale", but I do not view them as fundamentally equivalent actions.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#364Earlier quoted context omitted.
So I can use an open source LLM like Llama then?
You've pierced my completely precise, absolutely airtight choice of language about this situation as some sort of flaw in the greater point being made. Less glibly: a non-profit oriented LLM is just in a little different place on the scale, but doesn't fundamentally change my takeaway. However in this situation it makes it particularly egregious.
Progress is murdering their business model. They are trying to stop it; which makes sense, but let’s not pretend the trajectory here isn’t to ubiquitous LLMs everywhere within 3 years. We have to plan for that, assess the benefit for society from that.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#365Earlier quoted context omitted.
IIRC the distinction is that facts can't be copyrighted, a particular arrangement of facts can, particularly if something more subjective (analysis or opinion) is included. So I can write my own article about the sky being shown to appear blue much of the time, but I can't copy someone else's article about the same subject.
How would one even go about a phonebook-style mechanical listing of facts and occurrences? You’re listing an impossibility and then saying that garden-variety connective sentences somehow make it not a factual listing. Like yes if you copy a NYT article verbatim it’s like copying a phone book ads and all, and that’s infringement. But that’s not what a LLM does, NYT doesn’t like their content being used and summarized…
The traditional trick there is to include some small amount of fake data in the directory. You know someone has copied your collection of facts instead of compiling their own because it includes your fake facts. Mapmakers have used the method for at least as long as cartography has been part of our recorded history, see https://en.wikipedia.org/wiki/Trap_street for details. As noted in that page, the legal status of this, like many IP related issues, depends upon jurisdiction.
> But that’s not what a LLM does,
What does it do that means it is only summarizing factual information? While NYT effectively trying to claim copyright on facts is wrong, OpenAI claiming it can't reproduce copyrightable information while it can reproduce/summarize facts found within the same training set seems at best disingenuous.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#366If a human reads something, it goes into their brain, and it becomes an influence on future works they produce. This doesn't mean that 'copywrite' extends into my brain. A company can't copywrite what I'm thinking about. And what if I do try to paraphrase something from memory, from a few sources, and happen to spit out a very similar sentence from memory. Am I breaking the law? To go further. Since all knowledge is…
> Am I breaking the law? The intent of [US] copyright law is to promote new works of art (which can be derivative). So copyright did exactly what it is supposed to do in your analogy. Plus, you're human, which gives you special rights that software doesn't posses.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#367Earlier quoted context omitted.
Does reading news articles not benefit you? If not why continue reading?
Exactly. If I read a NYT article, and decide to invest in some company, then sell the stock, make a profit. Do I owe the NYT a percentage because I used knowledge I "read" from one of their articles? I "read", input, (into my brain neural net) where it mixed with other inputs.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#368Earlier quoted context omitted.
How would one even go about a phonebook-style mechanical listing of facts and occurrences? You’re listing an impossibility and then saying that garden-variety connective sentences somehow make it not a factual listing. Like yes if you copy a NYT article verbatim it’s like copying a phone book ads and all, and that’s infringement. But that’s not what a LLM does, NYT doesn’t like their content being used and summarized…
> How would one even go about a phonebook-style mechanical listing of facts and occurrences? The traditional trick there is to include some small amount of fake data in the directory. You know someone has copied your collection of facts instead of compiling their own because it includes your fake facts. Mapmakers have used the method for at least as long as cartography has been part of our recorded history, see https…
> Trap streets are not copyrightable under the federal law of the United States. In Nester's Map & Guide Corp. v. Hagstrom Map Co. (1992),[3][4] a United States federal court found that copyright traps are not themselves protectable by copyright. There, the court stated: "[t]o treat 'false' facts interspersed among actual facts and represented as actual facts as fiction would mean that no one could ever reproduce or copy actual facts without risk of reproducing a false fact and thereby violating a copyright ... If such were the law, information could never be reproduced or widely disseminated." (Id. at 733)
And yes the EU has the concept of “database rights” but notionally there is still supposed to be a creative step required in the selection or arrangement of records. So just a raw copy of the numbers in a telephone directory is theoretically not copyrightable, but a telephone book might be because of the creative/transformational step. It’s possible this might be such a low bar that it’s impossible to fail to clear, but, at least on paper you can’t copyright mere facts and figures either.
But either way it’s generally true that simple facts and figures are not protected and trap streets are a discredited and clumsy attempt to work around this.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#369Earlier quoted context omitted.
Where do you see that they won the case? Can you provide a source because the wikipedia article directly contradicts what you are saying...? I see they went to the Supreme Court who kicked it back to the Ninth who then re-affirmed their position that HiQ Labs was not in violation of the CFAA.
From [0] and [1], it seems it was a mixed ruling. I am actually not sure whether it's now legal to scrape, since the Court ruled against hiQ due to a breach of terms of service, but previously the Ninth Circuit Court affirmed its ruling against LinkedIn. [0] https://www.natlawreview.com/article/court-finds-hiq-breache... [1] https://www.natlawreview.com/article/hiq-and-linkedin-reach-...
So at a federal level, it seems relatively clear. The only uncertainty is on the state level.
Re: New York Times considers legal action against OpenAI as copyright tensions swirl
#370Earlier quoted context omitted.
Short of a police state, how would you enforce this? This has napster -> subscription spotify energy. But the only people happy about that are Spotify and people who found it distasteful to download music illegally. There just wasn’t a consumer-friendly option for a while, so the black market was the only market. So. The enforcement mechanism is what… a scary DMCA letter? (There will definitely be a stupid DCAIA in t…
The copyright holder gets a share of ownership in any AI model derived from its work, and thus a share of any resulting revenue.
In an (unrealizable) regime where all copyright holders are compensated, that would include picopennies for the discussion we’ve had!