Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

791–800 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#791

Earlier quoted context omitted.

I’d like a productive comment if you have it.

I'd've liked very much to give you one, but your blanket dismissal of what I consider to be a very important step for the field of AI ethics as the work of "Innovation Sponges" reveals that we fundamentally disagree about some pretty big underlying issues here. In the future, if you wish to invite productive comments and not curt dismissal, consider framing your concerns as potential risks rather than the cynical exp…

Fair enough and your criticism is appreciated.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#792

Earlier quoted context omitted.

I see, the narrative switched form “cat’s out of the bag” to “genie’s out of the bottle”. Regardless, no one wants to ban llms. We just want the theft to stop.

Copying is not theft. Stealing a thing leaves one less left Copying it makes one thing more; that’s what copying’s for.

Quite the hill to die on. But hear me on this: why doesnt openai allow training using its data? Or why doesnt microsoft train against windows’ source code (not that it would be of quality)? Or why can’t we just copy whatever movie and music we want through whatever protocol we want? Is it because your bosses know that copying without approval is theft?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#793
This looks a lot more convincing to me than the Copilot lawsuit or the Sarah Silverman one. This suit shows ChatGPT reciting large amounts of NYT articles - not just little snippets of code or the ability to answer questions about Silverman's book.

It feels like even if training on copyrighted data is fair use (and I think it should be), that wouldn't give you a pass on regurgitating that training data to anyone who asks.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#794
post #736

Earlier quoted context omitted.

Nobody is gonna cancel their NYT subscription for chatGPT 4.0. OpenAI will win.

Per my other comment here, https://news.ycombinator.com/item?id=38784723 , courts have previously ruled that whether people would cancel their NYT subscription is irrelevant to that test.

What exactly is the effect on the potential market? That's exactly why I don't think OpenAI will lose, why would a court side with the NYT?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#795

Earlier quoted context omitted.

Why using authored NYT articles is “stupid IP battles” and having to pay for the trained model with them is not stupid?

If the NYT wanted to charge OpenAI $20/mo to access their articles like any other user, that's fine with me. But they're not asking for that, they're suing them to stop it instead. That's why it's a stupid IP battle.

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#796
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

I'm not for or against anything at this point until someone gets their balls out and clearly defines what copyright infringement means in this context. If you give a bunch of books to a kid all by the same author and then pay that kid to write a book in a similar style and then I go on to sell that book...have I somehow infringed copyright? The kids book at best is likely to be a very convincing facsimile of the orig…

I think your kid analogy is flawed because it ignores the fact that you couldn't reasonably use said "kid" to rapidly produce thousands of works in the same style and then go on to use them to flood the market and drown out the original authors presence.

Try this with a real "kid" and you'll run into all kids of real-world constraints whereas flooding the world with derivative drivel using LLMs is something that's actually possible.

So yeah, stop using weak analogies, it's not helpful or intelligent.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#797

Earlier quoted context omitted.

> But it seems to me that, if anything, machines should be held to higher standard than people. If machines achieve sentience, does this still hold? Like, we have to license material for our sentient AI to learn from? They can't just watch a movie or read a book like a normal human could without having the ability to more easily have that material influence new derived works (unlike say Eragon, which is shamelessly S…

As long as machines needs to leech on human creativity those humans needs to be paid somehow. The human ecosystem works fine thanks to the limitations of humans. A machine that could copy things with no abandon however could easily disrupt this ecosystem resulting in less new things being created in total, it just leeches without paying anything back unlike humans. If we make a machine that is capable of being as cre…

> If we make a machine that is capable of being as creative as humans and train it to coexist in that ecosystem then it would be fine. But that is a very unlikely case, it is much easier to make a dumb bot that plagiarizes content than to make something as creative as a human.

I disagree that our own creativity doesn't work that way: nothing is very original, our current art is based on 100k years of building up from when cave man would scrawl simple art into the stone (which they copied from nature). We are built for plagiarism, and only gross plagiarism is seen as immoral. Or perhaps, we generalize over several different sources, diluting plagiarism with abstraction?

We are still in the early days of this tech, we will be having very different conversations about it even as soon as 5 years later.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#798
i dont know. the new york times is keeping those pages online and accessible. if a human can go check those pages, and take notes, the same human can write code that will go read those pages and produce notes, or data based on the contents. call is AI if you like, doesnt matter. the nyt has that content online, and accessible. and on internet, there is no difference between a human grabbing that data, or a machine.

if you put content on internet and accessible to humans, why do you want to now say to people that if it's a machine that does it, suddenly you do not agree ? i am free to write code or design a machine to go get that data, and do whatever i want with it (as long as i don't do something illegal like stealing content under copyright)

and i don't give a F about the "terms of use" those morons put online, because those have NO value. there is either a contract signed by two parties, or there is not. and content you put on internet, and accessible to everyone that sends you a GET, is like writing stuff on a page, and putting that page outside on the street.

we could use humans to go read all those pages, and create new content from it from the knowledge gained on those various subjects. machine are here to reproduce what humans can do, to free us time for more interesting things. those servers that send data back from a GET, it is the same request when it's done by me, a human, or a machine. and those morons did put that data there, accessible to all, so now to see them cry foul makes me laugh.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#799

Earlier quoted context omitted.

All I can ever think about with how ML models work is that they sound an awful lot like Data Laundering schemes. You can get basically-but-not-quite-exactly the copyrighted material that it was trained on. Saw this a lot with some earlier image models where you could type in an artists name and get their work back. The fact that AI models are having to put up guardrails to prevent that sort of use is a good sign that…

>You can get basically-but-not-quite-exactly the copyrighted material that it was trained on. You can do exactly the same with a human author or artist if you prompt them to. And if you decide to publish this material, you're the one liable for breach of copyright, not the person you instructed to create the material.

Not if that person is a trillion dollar corporation. If they're a business that's regularly stealing content and re-writing it for their customers that business is gonna go down. Sure, a customer or two may go down with them but the business that sells counterfeit works to spec is not gonna last long.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#800

Earlier quoted context omitted.

Yes and OpenAI sells its copies as a subscription, so that’s at least copyright infringement if not theft.

They're not copies, no matter how much you want them to be.

They are copied. If I can say something like “make a picture of xyz in the style of Greg rutkowski” and it does so, then it’s a copy. It’s not analogous to a human because a human cannot reproduce things like a machine can. And if someone did copy someone artwork and try to sell it, then yes that would be theft. The logic doesn’t change just because it’s a machine doing it.
Post reply on HN