Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

141–150 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#141
post #92
post #44

Earlier quoted context omitted.

> Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? Because it's false equivalence? ChatGPT isn't a human being. It's a product that is built upon data from other sources. The question is if this data is legal to scrape, which it is: Web scraping is legal, US appeals court reaffirms [ https://news.ycom…

ChatGPT isn't doing the scraping, humans are. And humans are using computers to both read the article and create content or to scrape it. So not it's not a false equivalence.

How is it not scraping? There's no other way to get all that data for training a model without scraping.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#143
post #58
post #26

Earlier quoted context omitted.

You're saying plagiarism isn't if one mostly swaps a couple of things in the expression of the content?

the definition of plagiarism is "the practice of taking someone else's work or ideas and passing them off as one's own"...chatGPT might infer from thousands or millions of different possible works or ideas before creating its own sentences so I don't think that meets the definition of plagiarism

In general no, but there is a problem in that ChatGPT may end up "regurgitating" large chunks of source material regardless, even if mechanistically that's not what it's trying to do. Similarly it's been recently reported that Stable Diffusion has effectively memorized some entire images it was trained on, and is capable of generating those as output.

I don't think the word-by-word statistical mechanism of ChatGPT would stand up as a copyright defense in court. It's the output that counts, not the means of getting there. It'd be like me copying some copyright work word-for-word then trying to claim "well, your honor, I was only using that for inspiration, I was using my full creative abilities to write what I did, so you can't blame me if it's a word-for-word copy".

I think OpenAI (or any company with the resources to train such a model in the first place) could fairly easily self-police and check that what they are generating isn't an exact (or almost exact) copy of something it was trained on. It's a bit like the app Shazam/similar recognizing a song from a short snippet - you just need to generate some type of "hash code" for each generated sentence (or whatever level of granularity makes sense) and compare it to a database of "hash codes" from the source material it was trained on.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#144
post #63
post #44

Earlier quoted context omitted.

> Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? Because it's false equivalence? ChatGPT isn't a human being. It's a product that is built upon data from other sources. The question is if this data is legal to scrape, which it is: Web scraping is legal, US appeals court reaffirms [ https://news.ycom…

> The question is if this data is legal to scrape ...it is? I didn't see that question raised in OP's text at all. What do legacy human legalities have to do with how AI will behave? > Because it's false equivalence? ChatGPT isn't a human being. Is this important? What is so special about human learning that it puts it in a morally distinct category from the learning that our successors will do? It sounds like OP is…

>Is this important?

Well yes, it's the whole crux of the matter. Laws govern human behaviour. As of 2023, only living beings have agency. If I shoot someone with a gun, the criminal is me and not the gun. Being a deterministic piece of silicon, a computer is perfectly equivalent. Sure, it is important to start a discussion of potential nonhuman sentience in the future, but these AI models are not unlike any previous software in legal issues. It's bizarre to me how many people are missing this.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#145

Anybody have an idea of what kind of hardware (and cost) you would need to train the model and to execute it ? Obviously storage is not a major factor here.

I seem to recall that the training cost for ChatGPT was in the tens of millions of dollars. Execution cost is on the order of ~$1 per interaction.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#146
post #11

Do you want companies to do this in private for private gain and not share it with you? Because making it illegal will just make it happen in greater secrecy.

This is essentially a defeatist argument flirting with supporting extortion, it seems to me. If you think that chatgpt is doing something wrong, this is arguing that you should allow the wrong to exist because there's nothing you can do about it.

In other areas of society where a bad thing cannot be stopped, we still use legislation to reduce the amount of it and mitigate some of the harm.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#147

Ah, our daily dose of a bunch of people with basically no understanding of copyright law or even the basic concepts of tort or common law jurisprudence make all sorts of silly anthropomorphic arguments about “how computers think”. Please, people, learn how to focus your thoughts. Go read up on copyright law in the United States. If you go into learning about copyright law trying to justify your own preconceived notio…

Absolutely this, well put. I guess enough people misunderstand AI models to the point of treating them like they are not software. I guess this validates Clarke's third law (Any sufficiently advanced technology is indistinguishable from magic).

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#148
post #141
post #92

Earlier quoted context omitted.

ChatGPT isn't doing the scraping, humans are. And humans are using computers to both read the article and create content or to scrape it. So not it's not a false equivalence.

How is it not scraping? There's no other way to get all that data for training a model without scraping.

It's scraping both when humans do it and when the ChatGPT team do it, but that wasn't the point the parent made. He made a moral/philosophical point which is what i responded to.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#149
post #98

Earlier quoted context omitted.

Yes but that's a technical issue. I took the parent as making a philosophical point and responded in that spirit.

Wouldn’t it be nice if the people on these forums were not ignorant of both philosophy or the legal system before diving into incoherent conversations about both at the same time where the main thrust is the emotions they have about these tools?

yup

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#150
post #5

This argument came up a bunch a while back. I settled on the opinion that while it's possible to buy summaries of books, I don't give a fart in a breeze where ChatGPT got it's data. E.g. Summary of How to Win Friends and Influence People: Effective Steps to Better Interpersonal Relationships by Book Lyte ChatGPT does more of a mashup with the learned data than humans need to, that'll do me.

The problem is this perspective is from copyright owners and not chatgpt users. It's fine if you don't care, but what matters is--do courts and lawmakers care. Today is probably the right time to get started on it for those types.
Post reply on HN