Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

701–710 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#701

Earlier quoted context omitted.

> It's ultimately a kind of lossy statistical compression scheme at some level. And on this subject, it seems worthwhile to note that compression has never freed anyone from copyright/piracy considerations before. If I record a movie with a cell phone at a worse quality, that doesn't change things. If a book is copied and stored in some gzipped format where I can only read a page at a time, or only read a random page…

If you watch a bunch of movies then go on to make your own movie based on influence from these movies, you are protected even if you have mentally compressed them into your own movie. At some point, you can learn, be influenced and be inspired from copyrighted material (not copyright infringement), and at some point you are just making a poor copy of the material (definitely copyright infringement). LLMs are probably…

But that's not what ChatGPT is doing, or is it? ChatGPT watches and records a bunch of movies, then stitches together its own movie using scenes and frames from the movies it recorded. AI will never reach the former case until it learns to operate a camera.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#702

Earlier quoted context omitted.

> right to use the MIT That right ended when he used it to break the law. It was also for use on MIT computers, not for remote access (which is why he decided to install the laptop, also knowing this was against his "right to use"). The "right to use" also included a warning that misuse could result in state and federal prosecutions. It was not some free for all. > and pull the JSTR information No, he did not have th…

Ok, so a lot you've written but it comes down to this. What law did he break? Neither MIT nor JSTOR raised issue with what Schwartz did. JSTOR even went out of their way to tell the FBI they did not want him prosecuted. Remember, again, with what he was charged. Wiretapping and intent to distribute. He wasn't charged with trespassing, breaking and entering, or anything else. Wiretapping and intent to distribute. > Hi…

> What law did he break?

You can read the indictment, which I already suggested you do.

> Remember, again, with what he was charged. Wiretapping and intent to distribute. He wasn't charged with trespassing, breaking and entering, or anything else. Wiretapping and intent to distribute.

He wasn't charged with wiretapping (not even sure that's a generic crime). He was charged with (two counts of) wire fraud (18 USC 1343), a huge difference. He also had 5 different charges of computer fraud (18 USC 1030(a)(4), (b) & 2), 5 counts of unlawfully obtaining information from a protected computer (18 USC 1030 (a)(2), (b), (c)(2)(B)(iii) & 2), and 1 count of recklessly damaging a protected computer (18 USC...).

He was not charged with "intent to distribute", and there's not such thing as a "wiretapping" charge. Did you ever once read the actual indictment, or did you just make all this up from internet forum posts?

If you're going to start with the phrase "Remember, again.." you should try to make up nonsense. Actually read what you're asking others to "remember" which you apparently never knew in the first place.

> you are confusing a "crime" with "misuse of a system"

Apparently you are (willfully?) ignorant of law.

> You violate copyright by transforming.

That's false too. Transformative use is one defense used to not infringe copyright. Carefully read up on the topic.

> ask it to implement Java's Array list and get several verbatim parts of the JDKs source code echoed back at you

Provide the prompt. Courts have ruled that code that is the naïve way to create a simple solution is not copyrighted on it's own, so if you have only a few disconnected snippets, that violates nothing. Can you make it reproduce an entire source file, comments, legalese at the top? I doubt it. To violate copyright one needs a certain amount (determined by trials) of the content.

You might also want to make sure you're not simply reading OpenJDK.

> 0, because he didn't distribute.

Please read. "How many could I get from what Schwartz downloaded?" does not mean he published it all before he was stopped. It means what he took.

That you seem unable to tell the difference between someone copying millions of PDF to distribute as-is, and the effort one must go to to possibly get a desired copyrighted snippet, shows either dishonestly or ignorance of relevant laws.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#703
post #389
post #358

Earlier quoted context omitted.

Would be fascinated to hear from someone inside on a throwaway, but my nearest experience is that corporate lawyers aren't stupid. If there's legally-murky secret data sauce, it's firewalled from being easily seen in its entirety by anyone not golden-handcuffed to the company. They may be able to train against it. They may be able to peek at portions of it. But no one is downloading-all.

Big corporations and corporate lawyers lose major lawsuits all the time.

That doesn't mean they don't spend lots of time thinking of ways not to lose them.

See: Google turning off retention on internal conversations to avoid creating anti-trust evidence

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#704
post #701

Earlier quoted context omitted.

If you watch a bunch of movies then go on to make your own movie based on influence from these movies, you are protected even if you have mentally compressed them into your own movie. At some point, you can learn, be influenced and be inspired from copyrighted material (not copyright infringement), and at some point you are just making a poor copy of the material (definitely copyright infringement). LLMs are probably…

But that's not what ChatGPT is doing, or is it? ChatGPT watches and records a bunch of movies, then stitches together its own movie using scenes and frames from the movies it recorded. AI will never reach the former case until it learns to operate a camera.

How do you not know this isn’t what we are doing in some more advanced form? Anyways, the comparisons will become more apt as the tech advances.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#705

Earlier quoted context omitted.

The court decided to focus on the tiny snippets Google displayed rather than the full text on their servers backing the search functionality. The court found significant that Google deliberately limited the snippet view so it couldn't be used as a replacement for purchasing the original book. The opinion is a relatively easy read, I highly recommend it if you're interested in the issue. It's also notable the court co…

I referred to the case to point out that Google practically got away with what was a gigantic violation of the spirit or principle here, in that Google gets to keep a copy of these millions of works for itself without ever having paid for them, regardless of what it made available to the public. As for "what would Google do with all these book copies anyway if they can't make it public?", that has now been answered m…

The case was, like any case, based on the specific facts.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#706
I feel like there is a way to get around this where you use as many materials (books, newspapers, crawling websites etc) to generate the LLM so it can be good at reasoning/next token generation but it can only use reference knowledge files to answer your question that a user uploads at the time of asking.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#707
post #600

Earlier quoted context omitted.

for what it's worth, i asked altman directly and he denied using libgen or books2, but also deferred to murati and her team on specifics. but the Q&A wasn't recorded and they haven't answered my follow-ups.

Really? Because the GPT-3 paper talks about "...two internet-based books corpora (Books1 and Books2)..." (see pages 8 and 9) - https://arxiv.org/pdf/2005.14165.pdf Unclear what that corpora might be, or if its the same books2 you are referring to.

My guess is that this poster meant books3, not books2.

books1 and books2 are OpenAI corpuses that have never (to my knowledge) had their content revealed.

books3 is public, developed outside of OpenAI and we know exactly what's in it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#708
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

The world you’re hoping for will put all AI tech only within the hands of the established top 10 media entities, who traditionally have never compensated fairly anyway. Sorry but if that’s the alternative to some writers feeling slighted, I’ll choose for the writers to be sad and the tech to be free.

“Feeling slighted” is a gross understatement of how a lack of compensation flowing to creators has shaped the internet and the wider world over the past 25 years. If we have a problem with the way top media companies compensate their creators, that is a separate issue - not a justification for layering another issue on top.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#709
post #504
post #486

Earlier quoted context omitted.

Furthering the S3 health data thought exercise: If OpenAI got their hands on an S3 bucket from Aetna (or any major insurer) with full and complete health records on every American, due to Aetna lacking security or leaking a S3 bucket, should OpenAI or any other LLM provider be allowed to use the data in its training even if they strip out patient names before feeding it into training? The difference between this ques…

> should [they] be allowed to use this data in training…? Unequivocally, yes. LLMs have proved themselves to be useful, at times, very useful, sometimes invaluable assistants who work in different ways than us. If sticking health data into a training set for some other AI could create another class of AI which can augment humanity, great!! Patient privacy and the law can f*k off. I’m all for the greater good.

Eliminating the right to patient privacy does not serve the greater good. People have enough distrust of the medical system already. I’m ambivalent to training on properly anonymized health data but, i reject out of hand the idea that OpenAI et al should have unfettered access to identifiable private conversations between me and my doctor for the nebulous goal of some future improvement on llm models.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#710

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

"probably the single most important development in human history" is the kind of hyperbole you'd only find here. Better than medicine, agriculture, electrification, or music? That point of view simply does not jive with what I see so far from AI. It has had little impact beyond filling the internet with low-effort content. I feel like the crypto evangelists never got off the hype train. They just picked a new destina…

[deleted]
Post reply on HN