Live data from Hacker News

OpenAI pleads it can't make money with o using copyrighted material for free

futurism.com

21–23 of 23 posts

Re: OpenAI pleads it can't make money with o using copyrighted material for free

#22
post #11

Earlier quoted context omitted.

I have whiplash from your first and last sentences. > Students should be able to learn from books, music, film. So should AI training models. An AI model is a thing. It is owned and fully controlled by some agent. A student is a sentient, thinking being. Both can be trained, only one can be educated. Treating the two as comparable is misleading and in my view, wrong.

We're in strange new times, but the equivalence of human cognition and synthetic will likely become mainstream and mundane in the coming years. Sci-fi has long had various "cyborg" type things as a plot element, but if you walk down the street in NYC today you'll pass thousands of people with pacemakers, artificial hips, insulin pumps, colostomy bags, and prosthetics. People who've had laser surgery on their eyes to…

You are preoccupied with semantics and romantic notions of blurred lines between people and software, rather than the actual reality of what a model is, and who tends to control it. The "people" training models are mostly massive business interests that exist to create profit.

Re: OpenAI pleads it can't make money with o using copyrighted material for free

#23

Earlier quoted context omitted.

Isn't this about generating output after all? I'm not sure if I get your distinction about "consumption". > Do the current copyright laws not already protect the authors and give them tools for takedowns and remuneration? That was also my point in the prior HN comment thread on the MS news submission that I mentioned. Good luck starting "fair use" copyright lawsuits against a myriad of auto-generated derivatives. Thi…

If the goal is to prevent companies from training on copyright material, then yes, it is about consuming the material, not generating it. The generation part comes from anecdotal incidents where some copyright material has been generated. - This is not the normal - This can be changed over time, there are also moderation techniques that can be used. - We already have remedies for those publishing or selling copyright…

I'm not a luddite.

And I don't think that my argument was as narrow as you make it out to be.

It's not required to exactly reproduce training material for an AI to output something that wouldn't stand a "fair use" trial.

"Summarize XY, but prefer different words" is already enough for a blog post. And the possibility to do that is not limited to inference-time input.

Copyright law is about humans, not machines. The problem is scale. You deflected this argument instead of addressing it.

And regarding training: you seem to anthropomorphize LLMs in a weird way.

LLMs can only generate content that is entirely derived from their training data.

That the derivation is close to a blackbox for humans does not elevate machines to humans.

The burden of proof about training materials is IMO with LLM companies, not with human creators.

Because companies know full-well that anything that's not an obvious exact reproduction will require humans starting lawsuits in order to claim a copyright violation.

You say:

> - We already have remedies for those publishing or selling copyrighted material already

And I say, with regard to AI, you seem to be intentionally misinterpreting my comment.

Post reply on HN