Live data from Hacker News

OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

futurism.com

21–30 of 85 posts

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#21
post #16

Earlier quoted context omitted.

How is Google search breaking copywrite?

Exact same way OpenAI does: by scraping data, ingesting it, processing it, incorporating it into its proprietary system, and using it to serve responses to queries. This is not to say that I thing that any of this is wrong. I think that if what Google or OpenAI do is illegal, then the law is wrong, not Google or OpenAI.

Is it the exact same though?

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#22
post #16

Earlier quoted context omitted.

How is Google search breaking copywrite?

Exact same way OpenAI does: by scraping data, ingesting it, processing it, incorporating it into its proprietary system, and using it to serve responses to queries. This is not to say that I thing that any of this is wrong. I think that if what Google or OpenAI do is illegal, then the law is wrong, not Google or OpenAI.

ChatGPT is a substitute to the original work, Google redirects to the original work.

(and yes,the synthesis on the top of the page is a problem, I agree)

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#23
post #11

The product requires crime? I feel like most products do not require crime. This is not a good sales pitch.

Either that, or copyright law is bad in its current form and LLM’s are yet an example of what exposes that. Even if copyright owners can’t point to how much damage, if any, they suffer from AI, it’s seen as wrong and bad. I think it’s getting boring to hear that story about copyright repeat itself. In most crimes, you need to be able point to a damage that was done to you. Also, while there are edge cases in some LLM…

More like, it's interesting that big tech companies can create extremely elaborate copyright assignment, metering and payout mechanisms when it's in their interest - right down to figuring out who owns 30 seconds of incidental radio music that plays in the background during someone's speedrun video.

But for other classes of user generated content, the problem is suddenly "impossible".

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#24
post #16

Earlier quoted context omitted.

Exact same way OpenAI does: by scraping data, ingesting it, processing it, incorporating it into its proprietary system, and using it to serve responses to queries. This is not to say that I thing that any of this is wrong. I think that if what Google or OpenAI do is illegal, then the law is wrong, not Google or OpenAI.

Is it the exact same though?

Sorry, do you have actual point, or are just trying to be pedantic? Strictly speaking, it’s not exact same, because no two different things are exact same, by definition. However, my point is that the same principle should apply to both.

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#26

So basically, we know China is never going to pay the publishers/content creators ( never ). If we hold our principles to OpenAI ( pay who you took from ), they will go bankrupt. So of course they are speaking in end-game language. To suggest the race is lost even before it starts is an incredible thing. How is it that we can theorize that the model would get better with more data, but we can't theorize that the busi…

You know, there's a creative third way which the US could approach if it had the cajones.

Allow OpenAI and other AI companies to use all data for training, but require that they pay it forward by charging royalties on profits beyond X amount of profit, where X is a number high enough to imply true AGI was reached.

The royalties could go into a fund that would be paid out like social security payments for every American starting when they were 18 years old. Companies could likewise request a one time deferred payment or something like that.

It's having your cake and eating it. Also helping ease some tensions around job loss.

Sadly, what we'll likely get is a bunch of tech leaders stumbling into wild riches, hoarding it, and then having it taken from them by force after they become complacent and drunk on power without the necessary understanding of human nature or history to see why they've brought it on themselves.

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#27

I think we just need to rethink copyright for language models. I'm okay just licensing 1 copy of a work to any LLM model throughout its various generations. Just don't pirate it if no special license is available, buying the ebook should suffice. It should be no different from a human buying a copy. The rule should only be that it does not leak the entire work.

I'm not OK with that, though... and here we have the nut of the problem. There is no agreement as to what's acceptable and what's not.

I personally think that the odds of me me being able to both publicly publish my words and code and be able to keep them out of training data is pretty close to zero. Since that's unacceptable to me, my only option is not to publish that stuff at all.

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#29
post #16

Earlier quoted context omitted.

How is Google search breaking copywrite?

Exact same way OpenAI does: by scraping data, ingesting it, processing it, incorporating it into its proprietary system, and using it to serve responses to queries. This is not to say that I thing that any of this is wrong. I think that if what Google or OpenAI do is illegal, then the law is wrong, not Google or OpenAI.

Last I checked Google is not buying or pirating books for Google Search they just grab free data that has been provided.

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#30

So basically, we know China is never going to pay the publishers/content creators ( never ). If we hold our principles to OpenAI ( pay who you took from ), they will go bankrupt. So of course they are speaking in end-game language. To suggest the race is lost even before it starts is an incredible thing. How is it that we can theorize that the model would get better with more data, but we can't theorize that the busi…

You know, there's a creative third way which the US could approach if it had the cajones. Allow OpenAI and other AI companies to use all data for training, but require that they pay it forward by charging royalties on profits beyond X amount of profit, where X is a number high enough to imply true AGI was reached. The royalties could go into a fund that would be paid out like social security payments for every Americ…

Not to be funny on purpose, but we are having discussions in America currently on if we should finance aid for poverty and the like. I love your idea though.
Post reply on HN