Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

831–840 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#831

Earlier quoted context omitted.

The rules we have now were made in the context of human brains doing the learning from copyrighted material, not machine learning models. The limitations on what most humans can memorize and reproduce verbatim are extraordinarily different from an LLM. I think it only makes sense to re-explore these topics from a legal point of view given we’ve introduced something totally new.

Human brains are still the main legal agents in play. LLMs are just a computer programs used by humans. Suppose I research for a book that I'm writing - it doesn't matter whether I type it on a Mac, PC, or typewriter. It doesn't matter if I use the internet or the library. It doesn't matter if I use an AI powered voice-to-text keyboard or an AI assistant. If I release a book that has a chapter which was blatantly cop…

> Human brains are still the main legal agents in play.

No, they're not. This is The New York Times (a corporation) vs OpenAI and Microsoft (two more corporations).

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#832
post #491

Earlier quoted context omitted.

Can I ask what industries with what application? I've seen lots of task like summarizing articles or producing text. The image and video work seems too rudimentary to be taken seriously. Is there something out there that seems like a killer application? I was amazed at the idea of the block chain but we never found a use for it outside of cryptocurrency. I see a similariy with AI hype.

Well front page of HN right now is an article about how AI aided in the development of a new antibiotic

That wasn't an LLM trained on copywritten material.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#833
post #375

Earlier quoted context omitted.

Can you imagine spending decades of your life, studying skin cancer, only to have some $20/month ChatGPT index your latest findings and spit out generically to some subpar researcher: "Here's how I would cure melanoma!" followed by your detailed findings. Zero mention of you. F-that. Attribution, as best they can, is the least OpenAI can do as a service to humanity. It's a nod to all content creators that they have b…

Can you imagine spending decades of your life studying antibiotics, only to have an AI graph neural network beat you to the punch by conceiving an entire new class of antibiotics (first in 60 years) and then getting published in Nature. https://www.nature.com/articles/d41586-023-03668-1

As you already know yet are being intentionally daft about: They didn't use an LLM trained on copywritten material. There's a canyon of difference between leveraging AI as a tool, and AI leveraging you as a tool.

LLMs have, to my knowledge, made zero significant novel scientific discoveries. Much like crypto, they're a failure of technology to meaningfully move humanity forward; their only accomplishment is to parrot and remix information they've been trained on, which does have some interesting applications that have made Microsoft billions of dollars over the past 12 months, but let's drop the whole "they're going to save humanity and must be protected at any cost" charade. They're not AGI, and because no one has even a mote of dust of a clue as to what it will take to make AGI, its not remotely tenable to assert that they're even a stepping stone toward it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#834
post #421

Earlier quoted context omitted.

Because the issue isn’t the intake, it’s the output, where your analogy breaks down. If you could clone the brain of someone who was “trained” on decades of NYT and could reproduce its information on demand at scale, we’d be discussing similar issues.

Your analogy doesn't make sense either. If we could clone the brain of someone I hardly think we'd be discussing their vast knowledge of something so insignificant as the NYT. I don't think we should care that much about an AI's vast knowledge of the NYT either or why it matters. If all these journalism companies don't want to provide the content for free they're perfectly capable of throwing the entire website behin…

Most of the NYT is behind a signin screen; the classic "you can read the first paragraph of the page but pay us to see more" thing.

There is significant evidence (220,000 pages worth) in their lawsuit that ChatGPT was trained on text beyond that paywall.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#835

Earlier quoted context omitted.

It doesn't matter what is legal. It matters what is right. Society is about balancing the needs of the individual vs the collective. I have a hard time equating individual rights with the NYT and I know my general views on scraping public data and who I was rooting for in the LinkedIn case.

When we're discussing litigation, it certainly matters what is legal.

And also - if what is legal isn't right, we live in a democracy and should change that.

Saying what's legal is irrelevant is an odd take.

I like living in a place with a rule of law.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#836

Earlier quoted context omitted.

Ohh. You think being owner of a company whose newspaper is read by hundreds of millions of people every day, doesn't put you in a position of power to control the society?

I think I have better things to do than parse vague innuendo like that.

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#837
post #830

Earlier quoted context omitted.

> Humans have different rights than computer software Fortunately, the computer isn't the one being sued. Instead it is the humans who use the computer. And those humans maintain their existing rights, even if they use a computer.

Maybe (though there exist plenty of examples to the contrary). However, the NYT isn't suing you , ChatGPT user; they're suing OpenAI.

Gotcha.

OpenAI is run by humans as well though.

So the same argument applies.

Those humans have fair use rights as well.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#838

Earlier quoted context omitted.

How does AI compete with journalism? AI doesn't do investigative reporting, AI can't even observe the world or send out reporters. Which part of journalism is AI going to impact most? Opinion pieces that contain no new information? Summarizing past events?

AI certainly isn’t a replacement for journalism, but that doesn’t mean journalism will continue to exist if no one pays for it. If everyone gets their news from chatGPT or the like there will be no investigative reporting. We’re already beginning to see this with most people reading the google/Facebook blurbs instead of clicking the link and giving ad money let alone paying.

> If everyone gets their news from chatGPT

But I've just explained that ChatGPT can't actually produce news articles. I can't ask ChatGPT what happened today, and if I could it would be because a journalist went out and told ChatGPT what happened.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#839

Earlier quoted context omitted.

People often get buried in the weeds about the purpose of copyright. Let us not forget that the only reason copyright laws exist is > To promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right to their respective writings and discoveries If copyright is starting to impede rather than promote progress, then it needs to change to remain constitutional.

The reason copyright promotes progress is that it incentives individuals and organizations to release works publicly, knowing their works are protected against unlawful copying. The end game when large content producers like The New York Times are squeezed due to copyright not being enforced is that they will become more draconian in their DRM measures. If you don't like paywalls now, watch out for what happens if a…

I doubt there’s much that technical controls can do to limit the spread of NYT content, their only real recourse is to try suing unauthorized distributors. You only need to copy something once for it to be free.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#840
post #709
post #504

Earlier quoted context omitted.

> should [they] be allowed to use this data in training…? Unequivocally, yes. LLMs have proved themselves to be useful, at times, very useful, sometimes invaluable assistants who work in different ways than us. If sticking health data into a training set for some other AI could create another class of AI which can augment humanity, great!! Patient privacy and the law can f*k off. I’m all for the greater good.

Eliminating the right to patient privacy does not serve the greater good. People have enough distrust of the medical system already. I’m ambivalent to training on properly anonymized health data but, i reject out of hand the idea that OpenAI et al should have unfettered access to identifiable private conversations between me and my doctor for the nebulous goal of some future improvement on llm models.

> unfettered access to identifiable private conversations

You misread the post I was responding to. They were suggesting health data with PII removed.

Second, LLMs have proved that AI which gets unlimited training data can provide breakthroughs in AI capabilities. But they are not the whole universe of AIs. Some other AI tool, distinct from LLMs, which ingests en masse as much health data as it can could provide health and human longevity outcomes which could outweigh an individual's right to privacy.

If transformers can benefit from scale, why not some other, existing or yet to be found, AI technology?

We should be supporting a Common Crawl for health records, digitizing old health records, and shaming/forcing hospitals, research labs, and clinics into submitting all their data for a future AI to wade into and understand.

Post reply on HN