Live data from Hacker News

OpenAI asks White House for relief from state AI rules

finance.yahoo.com

641–650 of 845 posts

Re: OpenAI asks White House for relief from state AI rules

#642

Earlier quoted context omitted.

They do have to pay that. But if it's not fair use, they'd need to negotiate a custom license on top of that, for every single thing they use.

Where do they have to pay that? Where have they paid for each artwork from DeviantArt, paheal, etc that they trained Stable Diffusion on? Where have they paid for each independent blog post that they trained ChatGPT on? Yes, they've made a few deals with specific companies that host a large amount of content. That's a far cry from paying a fair price for each copyrighted work they ingest. Nearly everything on the Int…

Also, openai only started making deals (and mostly with news publishers) after the NYT lawsuit.

https://www.npr.org/2025/01/14/nx-s1-5258952/new-york-times-...

They didn't even consider doing this before. They still, as far as I know, haven't paid a dime for any book, or art beyond stock photography.

Lawsuit is still ongoing, if openai loses it might spell doom for legal production and usage of LLMs as a whole. There isn't enough open, free data out there to make state of the art AI.

Re: OpenAI asks White House for relief from state AI rules

#643
post #275

Earlier quoted context omitted.

"IP" is a very new concept in our culture and completely absent in other cultures. It was invented to prevent verbatim reprints of books, but even so, the publishing industry existed for hundreds of years before then. It's been expanded greatly in the past 50 years. Acting like copyright is some natural law of the universe that LLMs are upending simply because they can learn from written texts is silly. If you want t…

> It was invented to prevent verbatim reprints of books It was also invented to keep the publishing houses under control and keep them from papering the land in anti-crown propaganda (like the stuff that fueled the civil war in England and got Charles I beheaded). Probably one of the biggest brewing fights will be whether the models are free to tell the truth or whether they'll be mouthpieces for the ruling class. As…

That's why I am a big proponent of local, open-weights computation. They can't shut down a non-compliant model if you're the one running it yourself.

Re: OpenAI asks White House for relief from state AI rules

#644
post #623

Copyrighted material includes works by authors from outside the US. By Berne convention, the exceptions which any country may introduce must not "conflict with a normal exploitation of the work" and "unreasonably prejudice the legitimate interests of the author". So if at least one French author does license their work for AI training, then any exception of this kind will harm their legitimate interests and rob them…

> then any exception of this kind will harm their legitimate interests Pray tell what legitimate interest of the author is harmed by LLM's training on that work? No one is publishing the authors book.

What I think the parent meant is the interest to sell license to others to train on their data.

Re: OpenAI asks White House for relief from state AI rules

#645

If a person can read copyrighted material and produce derivative works, why not an AI?

Copyright does not restrict consumption. It only restricts reproduction. To restrict consumption you need a patent.

> It only restricts reproduction

and distribution.

Re: OpenAI asks White House for relief from state AI rules

#646

Earlier quoted context omitted.

Copyright does not restrict consumption. It only restricts reproduction. To restrict consumption you need a patent.

Good then that LLMs don't reproduce content.

They produce derivative works, which is also an exclusive right of a copyright holder.

Re: OpenAI asks White House for relief from state AI rules

#648
post #636

I think an AI should be treated like a human. A human can consume copyright material (possibly after paying for it), but not reproduce it. I don't see any reason why the same can't be true for an AI.

The issue is so much about consumption of copyright material, but acquisition of that material. Like a real person, AI companies need to adhere to IP and license or purchase the materials that they wish to consume. If AI companies licensed all materials they acquired for training purposes, this would be a non-issue. OpenAI are looking for a free pass to break copyright law, and through that, also avoid any issues tha…

A real person wouldn't have to pay to read random blog, Reddit comments, StackOverflow answers or code on GitHub (many open source licenses do not imply license for training).

They might have to pay for books, or use a library.

Should these cases be treated differently? If so, it might lead to more closed internet with even more paywalls.

Re: OpenAI asks White House for relief from state AI rules

#649

Earlier quoted context omitted.

They should train a model on a clean dataset and copyright dataset, charge extra on the copyright model, and pay a royalty to copyright owners when their works are cited in a response.

The problem there is how are we defining "works are cited"? Also couldn't you just do the same thing done to spotify and make bot farms to generate millions of citations?

You can simply pay to everyone whose works you have used for training, every time a model processes a request.

Re: OpenAI asks White House for relief from state AI rules

#650
post #592

Wonder how much the addition of copyrighted material affects how smart the resulting model is. If it's even 20% better LLM makers could be forced out of the US into jurisdictions that allow use of copyrighted data. I suspect most LLM users will ~always choose the smartest model.

The jump from llama2 to llama3 had something to do with meta downloading every textbook ever published and using it as training data. The arguments by meta so far in that court case are absolutely terrible and I'm half expecting to see the world's first trillion dollar copyright infringement award.

Incorrect. Llama 1 trained on books3 dataset.
Post reply on HN