Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

81–90 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#81
chatgpt actually has some ideas about this

question: How could the people who generate used in an ai language model be paid for their work?

answer: There are several ways in which the people who generate content for an AI language model could be paid for their work:

    Royalty-based payment: Content creators could receive a percentage of the revenue generated from the use of their content in the AI language model.

    Token-based payment: If the AI language model is built on a blockchain, content creators could be paid in tokens that could be traded for cryptocurrency or fiat currency.

    Partnership with content publishers: The developers of the AI language model could partner with content publishers to compensate the creators of the scraped content.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#83
My feeling is that one of the four happen:

1) AI is open sourced and we adapt stably. Either everybody has the opportunity to be their own business, or there is UBI.

2) AI is open sourced but it is unfairly distributed. Only some people are suited to BTOB, and/or UBI is shit.

3) AI is not open sourced, the wealthy edge out mankind and a planet scale genocide occurs.

4) none of it matters because the looming war between the US & China explodes or climate change wipes us out in any meaningful capacity that could pursue AI.

Given the track record of our species, #1 feels like wishful thinking

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#84
post #16

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

You usually expect people to cite sources. Granted, that very often doesn't happen, and the amount of citing expected depends on the context. But ChatGPT just doesn't cite sources at all. I think there's a case to be made that they should.

The bing leak seemed to mention sources.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#85
Yes it absolutely is, but imo less so than what GitHub Copilot and various image generation companies are doing. My theory is that if AI turns out to be as disruptive as the current hype suggests, the conflict between those who feed the AI vs. those who profit from it might be the next big social rift.

Artists are already in full rebellion against this, as they should be, being nearly eclipsed by AI, except when it comes to inventing new styles and hand-crafting samples for the models to train on. These, I assume, are either scraped off the web, or signed away in unfair ToS of various online publishing platforms.

Since the damage individually is small (they took some code from me without attribution, ok) but collectively enormous, in my opinion it the role of government to step in and soften the blow if necessary.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#86
post #73
post #44

Earlier quoted context omitted.

> Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? Because it's false equivalence? ChatGPT isn't a human being. It's a product that is built upon data from other sources. The question is if this data is legal to scrape, which it is: Web scraping is legal, US appeals court reaffirms [ https://news.ycom…

Being allowed to scrape something does not absolve you of all intellectual property, copyright, moral, etc. issues arising from subsequent use of the scraped data.

Exactly, besides, the question isn’t about legality, it’s about what the law should be, I think. The question isn’t whether it’s legal, the question is whether we need to change the law in response to technology.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#87

Earlier quoted context omitted.

Do you think just maybe there is a diffence here because humans need money to survive, and maybe we should have compassion for humans who could hypothetically starve or freeze or suicide or whatever because they have no money? Or is it just silly to care about people like that?

That's got nothing to do with whether or not it is "fair" for a learning system to produce content after it has learned. That is, instead, one of the larger and vastly more important sociocultural issues that actually warrants attention, but never receives it in sufficient degree to address the problem, because, for example, we're arguing whether automated learning is "fair".

If "fairness" isn't worth figuring out for a society, why is our entire economic order nominally built ontop of such a virtue? How is this not the very thing we are talking about? People starve on the streets right now because any other arrangement of resources has been deemed "unfair." Do we not sign a contract for our labor or for our homes because of shared idea of fairness? Fairness is the ultimate thing we appeal to in our world, it is the only thing that can sustain the intense individuality of the modern world. Dont ambiguate it as a Nietzchean moral fairness here, we are talking about the pseudo-algorithmic fairness of a market which guarantees certain things if you trade enough of your resources.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#88
How could training an AI on the works of Shakespeare possibly be unfair to him? Or to any other long dead person? - I don't see any issues

How could training an AI on the works of someone who has already been paid for them be unfair? - Possibly because it effects their future marketability and income?

Current authors, artists, internet commenters, clearly have an interest in the results of their creative endeavors being used for gain that they won't benefit from. This is very similar to the extractive monopolies of YouTube and the rest of social media. Their profit at our expense.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#89
Maybe unfair is the wrong word. I think most agree that scraping, even at a massive scale -- is in itself fair. But is it sustainable?

Will LLMs drive interest/activity away from wikipedia.org? Will it put its own sources of high-quality ad-supported content -- wikihow.com, for example (though I can't be totally sure it scraped from there) -- out of business? Or is there an earth-shattering copyright suit against OpenAI in the works as we speak?

> Can this start breaking the ad-based model of the internet

Is the alternative that everything is behind some kind of paywall by default, to block scraping? Is that where we're heading?

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#90
All I know is that while this isn't a new issue, the likes of ChatGPT has brought it to a head and made it more urgent. I am seriously reconsidering whether or not I want my writings to be available on the internet at all. I object to many of the uses, including this, they can be put to, and not publishing them online appears to be the only control available.

For now, I have removed my existing works, both technical and creative, from the internet and won't be adding more while I try to work out what to do.

Post reply on HN