Live data from Hacker News

OpenAI: Copy, Steal, Paste

computerworld.com

51–60 of 79 posts

Re: OpenAI: Copy, Steal, Paste

#51

Technically it's not a copy. And nothing was stolen. It's a best fit curve among a series of datapoints. The datapoint is a copy, if the best fit curve never touches the datapoint then it's technically not a copy. The difficult part is the technicality here is legal and ethical from any standpoint. The high level ramifications are a bit unfair in the sense that yes the data is being used to create an AI that can repl…

Copyright laws do not support your argument. There have been many cases in music where the offending song was forced to pay because it was "close enough" to the curve but not touching it.

I think it's an apt analogy, though I disagree about the implication.

If I use ChatGPT to create a work, and that work is "close enough" to an existing copyrighted work, then it seems like I am guilty of copyright violation, not ChatGPT.

Re: OpenAI: Copy, Steal, Paste

#52
post #41

"Just. Pay. Me." Ok, how much? How would one determine all sources and their contribution weights for each produced completion from GPT?

Easy!

(1) Make training opt-out trivial and fast

(2) if OpenAI still wants to use material in training, the can approach author with an offer.

Something like Google is a win-win proposition: Google gets useful results, while publisher gets clicks and ad impressions. But being included into AI training is not like that, authors get nothing for being included in the AI training set. So they should be free to charge whatever amount of money they want, or opt out completely.

Re: OpenAI: Copy, Steal, Paste

#53

I really hate this article. They compared training ChatGPT to actual copy&paste sites without any supporting argument, complained about how much money OpenAI is making without paying them, and finished with blaming generative AI for the declining quality in Google search results.

Have you seen the actual court PDF? Take a look, the link is in the article. There are many examples of ChatGPT exactly returning NYT content.

Re: OpenAI: Copy, Steal, Paste

#54
post #23
post #2

Can't agree more. These AI systems have been built on the back of free labor for decades. We will look back on this and wonder how we essentially subsidized these behemoth corporations, then allowed them to extract subscription fees back from the very people it took the labor from. Don't get me wrong, there are genuine innovations in AI and ML. But, on the same token, you can't have ChatGPT without content.

> These AI systems have been built on the back of free labor for decades. Do you repay publishers for information you’ve summarized or new insights you’ve gained after consuming their content? How is it different when AI does the same thing?

You can't really compare machines with people when it comes to scale.

You can scale the volume of work for a machine in a way that would be impossible with people. No amount of Mechanical Turks could make content at the rate that Generative AI systems like ChatGPT, Dall-E, etc. can.

Treating content scraped from an original source without consent as fair use for training a machine somehow seems wrong. Especially since this AI system is then used to make money and that content is essential to its ability to do so, though without giving anything back to the original content creators.

Seems natural that any law regarding AI copyright should take these matters into account.

Re: OpenAI: Copy, Steal, Paste

#55
post #34
post #21

The (in my view) problem with the author's argument is that the first step he claims is happening, is not. Publicly available content gets read, as is the point of publicly publishing it. Then the user uses a computer program to make some statistics about the bit of content. Those bits of statistics about that specific work, on their own, cannot reproduce or recrate the specific work. Then those statistics are put in…

Forgive me for this one, but it comes from genuine curiosity and not snark. You are making assertions about how copyright law works, but you don't qualify it with either IANAL or any legal credentials. So I must ask: do you have a basis for these claims? I love participating in armchair analysis of the law, since in software we pretty much have no choice but to do so anyway, but my understanding has always been that…

The EFF's post about it at each of the three steps of obtaining, training, and generating an output image: https://www.eff.org/deeplinks/2023/04/how-we-think-about-cop...

It's an interesting read and makes a good case for why none of the steps are directly copyright infringement, even if you can prompt the output to be (and in that case the person doing the prompting should be the one at fault, same as someone drawing something infringing directly).

Re: OpenAI: Copy, Steal, Paste

#56
post #16

Hard disagree. The whole concept of copyright is unnatural and technology evolution is showing its limits. In the same way a writer reads books and use that knowledge to create more, an ai model can do the same.

Imagine getting downvoted for this on a forum called Hacker News. Embarrassing that we have become copyright zealots.

I think it comes with creating content yourself.

Once you make some, you kinda want to have some control over it.

Re: OpenAI: Copy, Steal, Paste

#57

I'm a bit dubious about all these complaints. My snarky side wants to respond to > By OpenAI's logic, any work you put online is fair game to be swiped and incorporated into the company's large language models. With > Any work you put online is fair game to be swiped and incorporated into a human's brain. It's tough because the ability to copy an LLM is feasible, and an LLM is much more likely to be able to reproduce…

LLMs aren't human brains. The comparison only makes sense if one considers ChatGPT to be a sentient lifeform…which it so clearly is not.

Re: OpenAI: Copy, Steal, Paste

#58

Can someone explain how adjusting weights based on viewing data is copyright infringement? I don't know if it's just not understood how these models work, or if it's just purposely misleading to try and cripple scary new tech. It seems like writers trying to complain it's copyright infringement to have other writers read their works for inspiration.

Viewing alone would be fine, but they're not doing just that, are they?

Re: OpenAI: Copy, Steal, Paste

#59
post #41

"Just. Pay. Me." Ok, how much? How would one determine all sources and their contribution weights for each produced completion from GPT?

Not my problem. (Says the copyright holders.)

Look, I publish (nearly) all my blog posts for _free_ on the web because I like doing that, and I _still_ don't want chatbots scraping my content to serve their algos (yes, I'm blocking them now via robots.txt) because the reason I put content out for free is partially because I want my name and my expertise out there. I'm part of a community of authors if you will. Call it selfish or self-serving if you will, but that's my right as the author.

Having stuff I've contributed to the internet get ingested into a model and regurgitated out such that my contributions are sidelined…well, that's just not acceptable, any more than a person taking my blog post, rewriting it with a handful of other brief sources, then republishing it under their name with no attribution back to me is acceptable.

Re: OpenAI: Copy, Steal, Paste

#60

Earlier quoted context omitted.

Copyright laws do not support your argument. There have been many cases in music where the offending song was forced to pay because it was "close enough" to the curve but not touching it.

I think it's an apt analogy, though I disagree about the implication. If I use ChatGPT to create a work, and that work is "close enough" to an existing copyrighted work, then it seems like I am guilty of copyright violation, not ChatGPT.

Or both: when downloading music, both the one downloading and the one uploading can see legal action.
Post reply on HN