Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

771–780 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#771

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

Why using authored NYT articles is “stupid IP battles” and having to pay for the trained model with them is not stupid?

If the NYT wanted to charge OpenAI $20/mo to access their articles like any other user, that's fine with me. But they're not asking for that, they're suing them to stop it instead. That's why it's a stupid IP battle.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#772
post #375

Earlier quoted context omitted.

Can you imagine spending decades of your life, studying skin cancer, only to have some $20/month ChatGPT index your latest findings and spit out generically to some subpar researcher: "Here's how I would cure melanoma!" followed by your detailed findings. Zero mention of you. F-that. Attribution, as best they can, is the least OpenAI can do as a service to humanity. It's a nod to all content creators that they have b…

If someone paid me to study cancer and I discovered a cure, I'd give it away with or without credit. Who cares? If someone takes my software and uses it, cool. If they credit me, cool. If they don't, oh well. I'd still code. Not everything needs to be ego driven. As long as the cancer researcher (and the future robots working alongside them) can make a living, I really don't think it matters whether they get credit o…

You sounds like you’re trying to be cool or karma farming ?

I have no idea who invented the CT scanner, Xray machines, the hyperdermic needle, etc. I don't really care.

Maybe you should care because those things didn’t fall out do the sky and someone sure as shit got paid to develop and build those things. You copy and pasted code is worth less, a CT scanner isn’t.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#773

Earlier quoted context omitted.

Except Gmail does not own the copyrights to the email. So in the context to this article and theme of the post, the owner of the data is king. I don’t think any court would rule Google owned a novel sent over Gmail, little alone the contents of more normative emails.

from the Google Terms of Service ( https://policies.google.com/privacy?hl=en-US ), makes me wonder who owns what, since users of Gmail agree to it. "We also collect the content you create, upload, or receive from others when using our services. This includes things like email you write and receive, photos and videos you save, docs and spreadsheets you create, and comments you make on YouTube videos."

Collect does not mean own.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#774

Earlier quoted context omitted.

Why don’t they train their AI on non-copyrighted material? It’s only fair for the copyright owners to want a share of the pie. I’d want one as well for my work.

>It’s only fair for the copyright owners to want a share of the pie. No it's not, it's pure greed. Everyone'd think it absurd if copyright holders dared to demand that any human who reads their publicly available text has to pay them a fee, but just because OpenAI are training a brain made of silicon instead of a brain made of carbon all the rent-seekers come out to try to take advantage.

You know the NYT has to fork out money to build the content right ?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#775
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

Same, to all those arguing in favour of Open AI, I have a question, do you steal books, movies, games ?

Do you illegally share them via torrents or even sell copies of these works ?

Because that is what’s going on here?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#776

Earlier quoted context omitted.

I'm not so sure about that. It seems to me that they're selling me a service. Just like I might pay for a subscription to Adobe Photoshop or pay per-render fees to a rendering farm. I could use Photoshop to reproduce a copyrighted work, and in some circumstances (i.e. personal use) that'd be fine. Or I could use Photoshop to reproduce a copyrighted work and try to sell it for profit, which would clearly not be fine.…

The difference here is that Adobe is selling a set of tools that can recreate copyrighted work from the ground up. The Mona Lisa being previously incorporated into their tools is not a foundational necessity for their paintbrush to brush digital paint. The same is not true for AI, which require copyrighted work be contained therein, in order for the tool part to function.

While I 100% agree, there is another angle to consider this from, in that ChatGPT replaces reading the NYT. ChatGPT competes with it in the delivery of information.

To add to your point though, a sufficiently advanced AI trained on licensed data could reproduce copywrited content from prompt alone. It's the next step that would cause infringement where someone does something withcthe output.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#777

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

> This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. Easy to grandstand when it is not your job on the line.

Is it? My job as a frontend dev is similarly threatened by OpenAI, maybe even more so than journalists'. The very company I usually like to pay to help with my work (Vercel) is in the process of using that same money to replace me with AI as we speak, lol (https://vercel.com/blog/announcing-v0-generative-ui). I'm not complaining. I think it's great progress, even if it'll make me obsolete soon.

I was a journalism student in college, long before ML became a threat, and even then it was a dying industry. I chose not to enter it because the prospects were so bleak. Then a few months ago I actually tried to get a journalism job locally, but never heard back. The former reporter there also left because the pay wasn't enough for the costs of living in this area, but that had nothing to do with OpenAI. It's just a really tough industry.

And even as a web dev, I knew it was only a matter of time before I became unnecessary. Whether it was Wordpress or SquareSpace or Skynet, it was bound to happen at some point. I'm going back to school now to try to enter another field altogether, in part because the writing is on the ~~wall~~ chatbox for us.

I don't think we as a society owe it to any profession to artificially keep it alive as it's historically been. We do it owe it to INDIVIDUALS -- fellow citizens/residents -- to provide them with some way forward, but I'd prefer that be reskilling and social support programs, welfare if nothing else, rather than using ancient copyright law to favor old dying industries over new ones that can actually have a much bigger impact.

In my eyes, the NYT is just another news outlet. A decent one, sure, but not anything substantially different than WaPo or the LA Times or whatever. How many Pulitzer winners have come and gone? https://en.wikipedia.org/wiki/Pulitzer_Prize_for_Breaking_Ne...

If we lost the NYT, it'd be a bit of nostalgia, but next week life would go on as usual. They're not even as specialized as, say, National Geographic or PopSci or The Information or 404 Media or The Center for Investigative Reporting, any of which would be harder to replace than another generic big news outlet.

AI, meanwhile, has the potential to be way bigger than even the Internet, IMO, and we should be devoting Manhattan Project-like resources to it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#778
post #736

Earlier quoted context omitted.

It's likely not. Search for "the four factors of fair use". While I think OpenAI will have decent arguments for 3 of the factors, they'll get killed on the fourth factor, "the effect of the use on the potential market", which is what this lawsuit is really about. If your "fair use" substantially negatively affects the market for the original source material, which I think is fairly clear in this case, the courts wont…

Nobody is gonna cancel their NYT subscription for chatGPT 4.0. OpenAI will win.

Per my other comment here, https://news.ycombinator.com/item?id=38784723, courts have previously ruled that whether people would cancel their NYT subscription is irrelevant to that test.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#779

Earlier quoted context omitted.

There's no obvious need to hold people / AI to same standards here, yet, even if compression in mental-models is exactly analogous to compression in machine-models. I guess we decided already that corporations are already "like" persons legally, but the jury is still out on AIs. Perhaps people should be allowed more leeway to make possibly-questionable derivative works, because they have lives to live, and genuine if…

> But it seems to me that, if anything, machines should be held to higher standard than people. If machines achieve sentience, does this still hold? Like, we have to license material for our sentient AI to learn from? They can't just watch a movie or read a book like a normal human could without having the ability to more easily have that material influence new derived works (unlike say Eragon, which is shamelessly S…

As long as machines needs to leech on human creativity those humans needs to be paid somehow. The human ecosystem works fine thanks to the limitations of humans. A machine that could copy things with no abandon however could easily disrupt this ecosystem resulting in less new things being created in total, it just leeches without paying anything back unlike humans.

If we make a machine that is capable of being as creative as humans and train it to coexist in that ecosystem then it would be fine. But that is a very unlikely case, it is much easier to make a dumb bot that plagiarizes content than to make something as creative as a human.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#780
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

I'm not for or against anything at this point until someone gets their balls out and clearly defines what copyright infringement means in this context. If you give a bunch of books to a kid all by the same author and then pay that kid to write a book in a similar style and then I go on to sell that book...have I somehow infringed copyright? The kids book at best is likely to be a very convincing facsimile of the orig…

I don't have a comment on your hypothetical, but this case seems to go far beyond that. If you read the actual filing at the bottom of the linked page, NYT provides examples where ChatGPT recited exact multi-paragraph sections of their articles and tried to pass it off as its own words. Plainly reproducing a work is pretty much the only situation where "is this copyright violation?" isn't really in flux. It's not dissimilar to selling PDFs of copywritten books.

If NYT were fully rellying on the argument that training a model in wordcraft using their materials is always copyright violation, or only had short quotes to point to, the philosophical debate you're trying to have would be more relevant.

Post reply on HN