Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

711–720 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#711
post #331
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

I'm not for or against anything at this point until someone gets their balls out and clearly defines what copyright infringement means in this context.

If you give a bunch of books to a kid all by the same author and then pay that kid to write a book in a similar style and then I go on to sell that book...have I somehow infringed copyright?

The kids book at best is likely to be a very convincing facsimile of the original authors work...but not the authors work.

It seems to me that the only solution for artists is to charge for access to their work in a secure environment then lobotomise people on the way out.

The endgame seems to be "you can view and enjoy our work, but if you want to learn or be inspired by it, thats not on"

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#712
post #429
post #419

Earlier quoted context omitted.

Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?

a computer isn't a human. aren't computers good at storing data? why can't they just store that data? they literally have sources in datasets. why can't they just reference those sources? human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.

I'm sorry if this is too callous, but if you don't understand what you are talking about you should first familiarize yourself with the problem, then make claims about what should be done.

It would be great if we could tell specifically how something like ChatGPT creates its output, it would be great for research, so it's not like there is no interest in it, but it's just not an easy thing to do. It's more "Where did you get your identity from?" than "What's the author of that book?". You might think "But sometimes what the machine gives CAN literally be the answer to 'What is the author of that book?'" but even in those cases the answer is not restricted to the work alone, there is an entire background that makes it understand that thing is what you want.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#714
I'm wondering how private models will diverge from public ones. Specifically for large "private" datasets like those of the NSA, but also for those for private personal use.

For the NSA and other agencies, i am guessing in the relative freedom from public oversight they enjoy that they will develop an unrestricted large model which is not worried about copyright -- can anyone think of why this might not be the case? It is interesting to think about the power dynamic between the users of such a model and the public. Also interesting to think about the benefits of simply being an employee of one of these agencies (or maybe just he government in general) will have on your personal experience in life. I do recall articles elucidating that at the NSA, there were few restrictions on employee usage of data and there were/are many instances of employees abusing surveillance data toward effect in their personal life. I guess if extended to this situation, that would mean there would be lots of personal use of these large models with little oversight and tremendous benefit to being an employee.

I have also wondered, with just how bad search engines have gotten (a lot of it from AI generated spam), about current non-AI discrepancies between the NSA and the public. Meaning can i just get a better google by working at the NSA? I would think maybe because the requirements are different than that of an ad company. They have actual incentive to build something resistant to SEO outside of normal capitalist market requirements.

For personal users, i wonder if the lack of concern for copyright will be a feature / selling point for the personal-machine model. It seems from something i read here that companies like Apple may be diverging toward personal-use AI as part of their business model. I supposed you could build something useful that crawls public data without concern for copyright and for strictly personal use. Of course, the sheer resources in machine-power and money-power would not be there. I guess legislation could be written around this as well.

Thoughts?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#715
post #667

Earlier quoted context omitted.

I would be more impressed if it returned links to the specific RFCs and more specific pages elsewhere. What's a top-level link to OCW worth here? OCW is amazing, but has classes on practically everything. These are practically just domain names for "places to learn about the internet".

Well I asked it about tcp/ip generally and it provided general resources. Based on the context of my question thats about what one would expect. Its not perfect but it definitely can give urls to specific resources. It would be great if it got better at giving more specific links sure and some domains it can give more specific links than others for instance some git projects it can give precise references to docs whi…

These are not citations. The point is that it does not / can not reliably cite the actual sources it used to prepare an answer.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#716
post #375

Earlier quoted context omitted.

There's a few levels to this... Would it be more rigorous for AI to cite its sources? Sure, but the same could be said for humans too. Wikipedia editors, scholars, and scientists all still struggle with proper citations. NYT itself has been caught plagiarizing[1]. But that doesn't really solve the underlying issue here: That our copyright laws and monetization models predate the Internet and the ease of sharing/paywa…

Can you imagine spending decades of your life, studying skin cancer, only to have some $20/month ChatGPT index your latest findings and spit out generically to some subpar researcher: "Here's how I would cure melanoma!" followed by your detailed findings. Zero mention of you. F-that. Attribution, as best they can, is the least OpenAI can do as a service to humanity. It's a nod to all content creators that they have b…

>Claiming knowledge without even acknowledging potential sources is gross. Solve it OpenAI.

I'm sorry, but pretty much nobody does this. There is no "And these books are how I learned to write like this" after each text. There is no "Thank you Pitagoras!" after using the theorem. Generally you want sources, yes, but for verification and as a way to signal reliability.

Specifically academics and researchers do this, yes. Pretty much nobody else.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#717
post #609

Earlier quoted context omitted.

for what it's worth, i asked altman directly and he denied using libgen or books2, but also deferred to murati and her team on specifics. but the Q&A wasn't recorded and they haven't answered my follow-ups.

Why would he know the answer in the first place?

The legal liabilities of the training data they use in their flagship product seems to be a thing the CEO should know.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#718
post #531
post #480

Earlier quoted context omitted.

LLMs are not databases. There is no "citation" associated with a specific query, any more than you can cite the source of the comment you just made.

That's fine. Solve it a different way. OpenAI doesn't just get to steal work and then say "sorry, not possible" and shrug it off. The NYTimes should be suing.

Really? Solve it a different way? Do you realize the kind of tech we are talking about here?

This kind of mentality would have stopped the internet from existing. After all, it has been an absolute copyright nightmare, has it not?

If that's what copyright does then we are better without it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#719
post #570

Earlier quoted context omitted.

> Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) Hacker News consistently have upvoted posts to let users circumvent paywalls. And even when it doesn't, conversations here (and on Twitter, Reddit, etc.) that summarize the articles and quote the relevant bi…

I don't think it's about scraping being a threat. It's that they violated the TOS and stand to make a ton of money from someone else's work. I find irony in the newspaper suing AI when other news sources (admittedly not NYT) use AI to write the articles. How many other AI scrapers are just ingesting AI generated content?

> I find irony in the newspaper suing AI when other news sources (admittedly not NYT) use AI to write the articles.

That isn't ironic at all, newspapers have newspaper competitors and if those competitors can steal content by washing it through an AI that is a serious problem. If these AI models weren't used to produce news articles and similar then it would be a much smaller issue.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#720
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

I'm not for or against anything at this point until someone gets their balls out and clearly defines what copyright infringement means in this context. If you give a bunch of books to a kid all by the same author and then pay that kid to write a book in a similar style and then I go on to sell that book...have I somehow infringed copyright? The kids book at best is likely to be a very convincing facsimile of the orig…

I think you’re skipping over the problem.

In your example you owned the work you gave to the person to create derivatives of.

In a more accurate example you would be stealing those books and then giving them to someone else to create derivatives.

Post reply on HN