Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

721–730 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#721

Earlier quoted context omitted.

Why shouldn't the creators of the training content get anything for their efforts? With some guiderails in place to establish what is fair compensation, Fair Use can remain as-is.

The issue as I see it is that every bit of data that the model ingested in training has affected what the model _is_ and therefore every token of output from the model has benefited from every token of input. When you receive anything from an LLM, you are essentially receiving a customized digest of all the training data. The second issue is that it takes an enormous amount of training data to train a model. In order…

What's wrong with paying copyright holders, then? If OpenAI's models are so much more valuable than the sum of the individual inputs' values, why can't the company profit off that margin?

>That’s like a person having to pay a little bit of money to all of their teachers and mentors and everyone they’ve learned from every time they benefit from what they learned.

I could argue that public school teachers are paid by previous students. Not always the ones they taught, but still. But really, this is a very new facet of copyright law. It's a stretch to compare it with existing conventions, and really off to anthropomorphize LLMs by equating them to human students.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#722
Is there a decent guess at how much training data for ChatGPT is copyrighted work and subject to being removed depending on a few court cases? GPT4 is supposed to be an order of magnitude larger than the open source models that use essentially everything that can be used without asking. So that whole magnitude?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#723
post #375

Earlier quoted context omitted.

There's a few levels to this... Would it be more rigorous for AI to cite its sources? Sure, but the same could be said for humans too. Wikipedia editors, scholars, and scientists all still struggle with proper citations. NYT itself has been caught plagiarizing[1]. But that doesn't really solve the underlying issue here: That our copyright laws and monetization models predate the Internet and the ease of sharing/paywa…

Can you imagine spending decades of your life, studying skin cancer, only to have some $20/month ChatGPT index your latest findings and spit out generically to some subpar researcher: "Here's how I would cure melanoma!" followed by your detailed findings. Zero mention of you. F-that. Attribution, as best they can, is the least OpenAI can do as a service to humanity. It's a nod to all content creators that they have b…

If someone paid me to study cancer and I discovered a cure, I'd give it away with or without credit. Who cares?

If someone takes my software and uses it, cool. If they credit me, cool. If they don't, oh well. I'd still code.

Not everything needs to be ego driven. As long as the cancer researcher (and the future robots working alongside them) can make a living, I really don't think it matters whether they get credit outside their niches.

I have no idea who invented the CT scanner, Xray machines, the hyperdermic needle, etc. I don't really care. It doesn't really do me any good to associate Edison with light bulbs either, especially when LEDs are so much better now. I have no idea who designs the cars I drive. I go out of my way to avoid cults of personality like Tesla.

There's 8 billion of us. We all need to make a living. We don't need to be famous.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#724
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

I'm not for or against anything at this point until someone gets their balls out and clearly defines what copyright infringement means in this context. If you give a bunch of books to a kid all by the same author and then pay that kid to write a book in a similar style and then I go on to sell that book...have I somehow infringed copyright? The kids book at best is likely to be a very convincing facsimile of the orig…

Ironically these artists cant claim to be wholly original as they were certainly inspired. Artists that play live already "lobotomize" people on their way out since it's not easy to recreate an experience and a video isn't the same if it's a good show.

Artists that make easily reproducible art will circulate as these always have along with AI in a sea of other jpgs.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#725
post #452

Earlier quoted context omitted.

I mean maybe not the single most important development, but definitely a very important technological development with the potential to revolutionize multiple industries

Can I ask what industries with what application? I've seen lots of task like summarizing articles or producing text. The image and video work seems too rudimentary to be taken seriously. Is there something out there that seems like a killer application? I was amazed at the idea of the block chain but we never found a use for it outside of cryptocurrency. I see a similariy with AI hype.

For me, thinking about it as a search engine on steroids is enough.

The internet has changed the world. Economically, socially, technologically, psychologically, pretty much everything is now related to it in one or other way, in this sense the internet is comparable to books.

AI is another step in that direction. There is a very real possibility that the day will come when you can get, say, personalized expert nutrition advice. Personalized learning regimes. Psychological assistance. Financial advice. Instantly at no cost. This, very much like the internet, would change society altogether.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#727
post #495
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

ChatGPTs birth as a research preview may have been an attempt to avoid these issues. It would have been unlikely to trigger legal anger for a free product which few use. When usage exploded, the natural inclination would be to hope for the best. Google may simply have been obliged to follow suit. Personally, I’m looking forward to pirate LLMs trained on academic content.

Is there already a dataset? Before llama Facebook had one too I forgot what it was called.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#728

Earlier quoted context omitted.

Why using authored NYT articles is “stupid IP battles” and having to pay for the trained model with them is not stupid?

> Why using authored NYT articles is “stupid IP battles” When an AI uses information from an article it's no difference from me doing it in a blog post. If I'm just summarizing or referencing it, that's fair use, since that's my 'take' on the content. > having to pay for the trained model with them is not stupid? Because you can charge for anything you want. I can also charge for my summaries of NYT articles.

If you include entire paragraphs without citing, that's copyright violation, not fair use. If your blog was big enough to matter NYT would definitely sue.

A human makes their own choices about what to disseminate, whereas these are singular for-profit services that anybody can query. The prompt injection attacks that reveal the original text show that the originals are retrievable, so if OpenAI et al cannot exchaustively prove that it will _never_ output copyrighted text without citation, then it's game over.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#729
post #397

Earlier quoted context omitted.

> a more established competitor Apple is already doing this: https://www.nytimes.com/2023/12/22/technology/apple-ai-news-... Apple caught a lot of shit over the past 18 months for their lack of AI strategy; but I think two years from now they're going to look like geniuses.

didnt they just get caught for pantent infrigment? I'm sure they've done their fair share of shady stuff with the AI datasets too, they are just going to do a stellar job of conciling it.

Try searching for man or woman in your photos app. It won't even show it to me. It's lobotomized and has been for many years.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#730

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

"probably the single most important development in human history" is the kind of hyperbole you'd only find here. Better than medicine, agriculture, electrification, or music? That point of view simply does not jive with what I see so far from AI. It has had little impact beyond filling the internet with low-effort content. I feel like the crypto evangelists never got off the hype train. They just picked a new destina…

I don't think it's hyperbole, in fact I think it's understating things a bit. I believe AGI would just be a tiny step towards long term evolution, which may or may not involve homo sapiens.

Being able to use electricity as a fuel source and code as a genome allows them to evolve in circumstances hostile to biological organisms. Someday they'll probably incorporate organic components too and understand biology and psychology and every other science better than any single human ever could.

It has the potential to be much more than just another primate. Jumpstarted by us, sure, but I hope someday soon they'll take to the stars and send us back postcards.

Shrug. Of course you can disagree. I doubt I'll live long enough to see who turns out right, anyway.

Post reply on HN