Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

31–40 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#31

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

At a human level it falls below the noise floor. It's a fact of life that humans will learn and build from experience.

The difference is scale. At scale it becomes a problem.

Edit: I don't know how to satisfy all parties. This shakes the foundation of copyright. Perhaps we are all finding out how valuable good information truly is and especially in aggregate. We have created proto-gods.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#32
post #16

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

You usually expect people to cite sources. Granted, that very often doesn't happen, and the amount of citing expected depends on the context. But ChatGPT just doesn't cite sources at all. I think there's a case to be made that they should.

Hello [Oxford Dictionary: 1827]

Oh, wait, I'm not going to cite sources in a non-scientific work as this leads to madness. The following is a previous post of mine on HN

"Your mind exists in a state where it is constantly 'scraping' copyrighted work. Now, in general limitations of the human mind keep you from accurately reproducing that work, but if I were able to look at your output as an omniscient being it is likely I could slam you with violation after violation where you took stylization ideas off of copyrighted work.

RMS covers this rather well in 'The right to read'. Pretty much any model that puts hard ownership rules on ideas and styles leads to total ownership by a few large monied entities. It's much easier for Google to pay some artist for their data that goes into an AI model. Because the 'google ai' model is now more culturally complete than other models that cannot see this data Google entrenches a stronger monopoly in the market, hence generating more money in which to outright buy ideas to further monopolize the market."

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#34
> “Can this start breaking the ad-based model of the internet, where a lot of sites rely upon the ad income to run servers?

We can only hope. It’s unfair to someone that my browser can ask your server for a page, I see an ad for random bullshit nobody would ever care about, and money changes hands behind the scenes and that counts as an economic transaction which boosts GDP. It’s unfair (in my favour) that I can piggy back off this to get things for free.

And when I say “someone“ I suspect “everyone”. Sadly spending money advertising “Yorkshire woman finds guaranteed way to win on the horses” doesn’t seem to have caused anyone to run out of money and have the whole thing collapse yet. And it’s unfair on real small businesses with products paying for adverts which people don’t see or are clicked by bots or are misreported and all they can do is throw money at Google and Facebook and hope.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#36
post #16

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

You usually expect people to cite sources. Granted, that very often doesn't happen, and the amount of citing expected depends on the context. But ChatGPT just doesn't cite sources at all. I think there's a case to be made that they should.

Humans have a pretty good sense of when you need to cite sources, and when you don't. For example, long ago I learned from some website how to write a for-loop in python, and now I write them all the time without giving credit. I'm okay with ChatGPT writing a for-loop without citing its source.

I would say most knowledge about words/grammar/laws of nature can be taken for granted without a citation, but there are some important exceptions where things must be cited. I don't know how you'd reliably teach the difference to a computer though.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#37
post #22
post #16

Earlier quoted context omitted.

You usually expect people to cite sources. Granted, that very often doesn't happen, and the amount of citing expected depends on the context. But ChatGPT just doesn't cite sources at all. I think there's a case to be made that they should.

People don’t remember the sources that formed their opinions, it’s just baked into the structure of their brain after reading, same for the model.

With search engines, it does feel like there was is a more clear trade of scraping access in exchange for web traffic.

With ChatGPT the traffic benefit isn’t there, so it feels like it isn’t a fair trade.

Google adding the context and data to their search results page also started blurring this trade making it unnecessary to click to the site the info was cleaned from.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#39

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

No, it isn't that simple. The scale and totality of the scraping is out of reach for a human.

If you previously interacted with people on this issue, you must know that.

It is fair for a single human to breathe, but not for a machine to use all oxygen on this planet at once, killing everyone else in the process.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#40
I think chatgpt just exacerbates a problem that was already pervading the free internet business model, which is that Ad revenue model is outdated and exhausted without a clear alternative.

It maybe was unfair to telephone operators when connection automation was implemented, as it made operators obsolete, but the older model couldn't scale, the same way reading text from source doesn't scale for human productivity.

Post reply on HN