Live data from Hacker News

OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

futurism.com

11–20 of 85 posts

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#11

The product requires crime? I feel like most products do not require crime. This is not a good sales pitch.

Either that, or copyright law is bad in its current form and LLM’s are yet an example of what exposes that.

Even if copyright owners can’t point to how much damage, if any, they suffer from AI, it’s seen as wrong and bad. I think it’s getting boring to hear that story about copyright repeat itself. In most crimes, you need to be able point to a damage that was done to you.

Also, while there are edge cases in some LLM’s where you can make them spew some verbatim training material, often through jailbreaks or whatnot, an LLM is a destructive process involving ”fuzzy logic” where the content is generally not perfectly memorized, and seems no more of a threat to copyright than recording broadcasts onto cassette tapes or VHS were back in the day. You’d be insane to use that stuff as a source of truth on par with the original article etc.

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#12
post #7

Earlier quoted context omitted.

By this logic Google Search couldn't exist. Except that Google won those cases.

How is Google search breaking copywrite?

The text preview underneath the search result is one thing I remember being contentious (some news websites in France took Google to court and won if I recall correctly)

For mostly the same reasons people are against AI. If you read that text, sometimes there’s no reason to visit the website, which ‘deprives’ website owners and the content creators of ad revenue they would have gotten if Google hadn’t copied the text from their website.

After all, the news is the same regardless of if it’s written on the Google result preview or on the news website itself.

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#14
So basically, we know China is never going to pay the publishers/content creators (never). If we hold our principles to OpenAI (pay who you took from), they will go bankrupt. So of course they are speaking in end-game language. To suggest the race is lost even before it starts is an incredible thing.

How is it that we can theorize that the model would get better with more data, but we can't theorize that the business model would need to get bigger (pay the content creators) to train the model? Shoot first and ask questions later (or rather, BEG later).

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#15
What I don't understand is why this is always presented as a "race" that "we" have to win or else. It's just such a strange framing to me and every time I see it, it's presented as some sort of self-evident truth, but I don't think it's self-evident at all.

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#16
post #7

Earlier quoted context omitted.

By this logic Google Search couldn't exist. Except that Google won those cases.

How is Google search breaking copywrite?

Exact same way OpenAI does: by scraping data, ingesting it, processing it, incorporating it into its proprietary system, and using it to serve responses to queries.

This is not to say that I thing that any of this is wrong. I think that if what Google or OpenAI do is illegal, then the law is wrong, not Google or OpenAI.

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#17
Yes, well, in a way they're right and I suspect everyone here knows it no matter how and mighty they might want act when commenting. When foreign (here 'Chinese') competition just ignores copyright laws while 'western' companies have to abide by them for every piece of data they use to train their models the former will have a clear advantage over the latter. This also happens to be how the USA acted in the 1800s [1]:

the United States declined an invitation to a pivotal conference in Berne in 1883, and did not sign the 1886 agreement of the Berne Convention which accorded national treatment to copyright holders. Moreover, until 1891 American statutes explicitly denied copyrights to citizens of other countries and the United States was notorious in the international sphere as a significant contributor to the "piracy" of foreign literary products. It has been claimed that American companies for the most part "indiscriminately reprinted books by foreign authors without even the pretence of acknowledgement" (Feather, 1994, 154). The tendency to freely reprint foreign works was encouraged by the existence of tariffs on imported books that ranged as high as 25 percent (see Dozer, 1949).

[1] http://socialsciences.scielo.org/scielo.php?script=sci_artte...

Re: OpenAI Says It's "Over" If It Can't Steal All Your Copyrighted Work

#18
Sorry but it is actually a huge problem for the US if the DeepSeek models are able to train on sorta-illegal dumps of scientific papers and US models aren't. The ones that are paywalled by scientific journals.

Everyone WILL start using hosted frontier Chinese models if they are demonstrably better at answering scientific questions than ChatGPT, sending essentially all US research questions into a Chinese data dump. This is even worse than the national security catastrophe that is TikTok (even aside from the EVEN BIGGER issue that China will have models that are staggeringly better than those in the US, because they are up to date on the science).

I understand the reflexivity against AI companies "stealing content" but we need to stay competitive and figure out the financial compensation later. This is not a case where our unbelievably generous copyright laws should take precedence over US competitiveness.

Post reply on HN