Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

871–880 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#871
post #816
post #459

Earlier quoted context omitted.

Except it doesn’t memorize text. It generates text that is statistically likely. Generating a citation that is statistically likely wouldn’t really help the problem.

So it's just bullshit then.

It's literally how our meat bag brains work pretty much.

Anything like word association games are basically the same exercise, but with humans and hell, I bet I could play a word association game with an LLM, too.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#872

Earlier quoted context omitted.

The model weights clearly encode certain full passages of text, otherwise it would be virtually impossible for the network to produce verbatim copies of text. The format is something very vaguely like "the most likely token after "call" is "me"; the most likely token after "call me" is "Ishmael". It's ultimately a kind of lossy statistical compression scheme at some level.

> It's ultimately a kind of lossy statistical compression scheme at some level. And on this subject, it seems worthwhile to note that compression has never freed anyone from copyright/piracy considerations before. If I record a movie with a cell phone at a worse quality, that doesn't change things. If a book is copied and stored in some gzipped format where I can only read a page at a time, or only read a random page…

Is it still compression if I read Tolkien and reference similar or exact concepts when writing my own works?

Having a magical ring in my book after I've read lord of the rings, is that copyright?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#873
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

> they probably wouldn’t exist and the generative AI revolution may never have happened if they put the horse before the cart

Maybe, but I find the "It's ok to break the law because otherwise I can't do what I want" narrative a little offputting.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#874
I've just tried to get gpt35 to spit out an NYT article verbatim in all sorts of ways and it just can't.

It will compile an article that "looks like" NYT's (or any other news site) but none of the paragraphs were a match for any of their articles that I could find.

I'm really curious to see what evidence they have for the case beyond "it can claim to be NYT and write an article composed of all sorts of bullshit from every corner of the Web".

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#875
post #845

Earlier quoted context omitted.

>All that aside, I tend to agree with the hypothesis that LLMs are a fad that will mostly pass. For professionals, it is really hard to get past hallucinations and the lack of citations. For writers maybe, but absolutely not for programmers, it's incredibly useful. I don't think anyone who's used GPT4 to improve their coding productivity would consider it a fad.

Copilot has been way more useful to me than GPT4. When I describe a complex problem where I want multiple solutions to compare, GPT4 is useless to me. The responses are almost always completely wrong or ignore half of the details I’ve written in the prompt. Or I have to write them with already a response in mind, which kinda defeats why I would use it in the first place. Copilot provides useful autocompletes maybe… 3…

> When I describe a complex problem where I want multiple solutions to compare, GPT4 is useless to me

FWIW i don’t try to use it for this. mostly i use it to automate writing code for tasks that are well specified, often transformations from one format to another. so yes, with a solution in mind. it mostly just saves typing, which is a minority of the work, but it is a useful time saver

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#876
post #678

Earlier quoted context omitted.

I don’t view LLMs as a fad. It’s like drummers and drum machines. Machines and drummers co-exist really well. I think drum machines, among other things, made drummers better.

It mainly made mediocre drummers sound better to the untrained ear.

It allowed people to see the difference between drum machines and humans. Drummers could practice to sound more like the ‘perfect’ machines, but more importantly the best drummers learned how to differentiate themselves from machines. The best drummers actually became more human. Listen and look at Nate Smith - this guy plays with timing and feel and audience reactions in ways that machines cannot. Sometimes tools let humans expand their creativity in ways previously unheard of. Just like the LLMs are doing right now.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#877
post #840
post #709

Earlier quoted context omitted.

Eliminating the right to patient privacy does not serve the greater good. People have enough distrust of the medical system already. I’m ambivalent to training on properly anonymized health data but, i reject out of hand the idea that OpenAI et al should have unfettered access to identifiable private conversations between me and my doctor for the nebulous goal of some future improvement on llm models.

> unfettered access to identifiable private conversations You misread the post I was responding to. They were suggesting health data with PII removed. Second, LLMs have proved that AI which gets unlimited training data can provide breakthroughs in AI capabilities. But they are not the whole universe of AIs. Some other AI tool, distinct from LLMs, which ingests en masse as much health data as it can could provide heal…

> Furthering the S3 health data thought exercise: If OpenAI got their hands on an S3 bucket from Aetna (or any major insurer) with full and complete health records on every American, due to Aetna lacking security or leaking a S3 bucket, should OpenAI or any other LLM provider be allowed to use the data in its training even if they strip out patient names before feeding it into training?

To me this says that openai would have access to ill-gotten raw patient data and would do the PII stripping themselves.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#878

If models are training on NYT content the future of AI is horrific.

We've banned this account for posting unsubstantive and/or flamebait comments and using HN primarily for ideological battle. That's not what this site is for.

If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#879

Earlier quoted context omitted.

> So at first ChatGPT will copy journalists. You still literally have not explained how this works. ChatGPT could write a news article, but it's not going to actively discover new social phenomena or interview people on the street. Niche journalism will continue having demand for the sole reason that AI can't reliably surface new and interesting content. So... again, how does a pre-trained transformer model scoop a j…

AI labs are working on and largely already have generative ai that can be actively updated. The generative ai scoops real journalists stories by watching their feed. This isn’t very different from the current status quo, it’s just a continuation of an already shitty situation for news organizations. If their revenue decreases even more than it already has they will cease to exist. Niche journalism barely has any dema…

Updated how? By what? Who is going out and investigating the world to write about? An AI does not have LEGS it can not go outside and go talk to someone and interview them, it can't attend a press conference without human assistance.

You have not at all explained how an AI is going to somehow write a news post about something that has just happened.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#880
post #367

Earlier quoted context omitted.

The law on this does not currently exist. It is in the process of being created by the courts and legistatures. I personally think that giving copyright holders control over who is legally allowed to view a work that has been made publicly available is a huge step in the wrong direction. One of those reasons is open source, but really that argument applies just as well to making sure that smaller companies have a cha…

It does exist, and you'd be glad to know that it's going in the pro-AI/training direction: https://www.reedsmith.com/en/perspectives/ai-in-entertainmen...

> It does exist, and you'd be glad to know that it's going in the pro-AI/training direction

Certainly not in the US. From the article you linked "In the United States, in the absence of a TDM exception, AI companies contend that inclusion of copyrighted materials in training sets constitute fair use eg not copyright infringement, which position remains to be evaluated by the courts."

Fair use is a defense against copyright infringement, but the whole question in the first place is whether generative AI training falls under fair use, and this case looks to be the biggest test of that (among others filed relatively recently).

Post reply on HN