Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?
91–100 of 194 posts
Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?
#92Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…
> Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? Because it's false equivalence? ChatGPT isn't a human being. It's a product that is built upon data from other sources. The question is if this data is legal to scrape, which it is: Web scraping is legal, US appeals court reaffirms [ https://news.ycom…
So not it's not a false equivalence.
Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?
#93Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?
#94It’s really just building a better model.
Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?
#95Earlier quoted context omitted.
> Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? Because it's false equivalence? ChatGPT isn't a human being. It's a product that is built upon data from other sources. The question is if this data is legal to scrape, which it is: Web scraping is legal, US appeals court reaffirms [ https://news.ycom…
ChatGPT isn't doing the scraping, humans are. And humans are using computers to both read the article and create content or to scrape it. So not it's not a false equivalence.
> Web scraping is legal, US appeals court reaffirms
First, the case is not closed. [0]
Second, to draw an analogy, you can use scraping in the same way you can use a computer: for legal purposes. That is, you cannot use scraping to violate copyright, just as you cannot use a computer to violate copyright.
The following being my conjecture (IANAL), there is fair use and there is copyright violation, and scraping can be used for either—it does not automatically make you a criminal, but neither is it automatically OK. If what you do is demonstrably fair use presumably you’d be fine; but OpenAI with its products cannot prove fair use in principle (and arguably the use stops being fair already at the point where it compiles works with intent to profit).
Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?
#96Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…
> Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? Because it's false equivalence? ChatGPT isn't a human being. It's a product that is built upon data from other sources. The question is if this data is legal to scrape, which it is: Web scraping is legal, US appeals court reaffirms [ https://news.ycom…
To be fair, so are you.
Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?
#97Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…
Does anyone actually find these arguments persuasive? There is really no reason to believe that what chatGPT or stable diffusion does is anything like what "your brain" does--except in the most superficial, inconsequential way. Second, try applying this logic to literally anything else and you'll see why it's absurd: "You can't ban cars from driving on sidewalks! If it's acceptable for people to walk on sidewalks, th…
I also agree it's not the only argument and ultimate proof.
I don't at, this point, have an answer. I'm sure this miraculous new technology will survive the luddite attacks, but there will probably be some tense moments, and some jurisdictions will choose to be left behind.
Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?
#98Earlier quoted context omitted.
ChatGPT isn't doing the scraping, humans are. And humans are using computers to both read the article and create content or to scrape it. So not it's not a false equivalence.
There’s a reason scraping is a legally grey area. > Web scraping is legal, US appeals court reaffirms First, the case is not closed. [0] Second, to draw an analogy, you can use scraping in the same way you can use a computer: for legal purposes. That is, you cannot use scraping to violate copyright, just as you cannot use a computer to violate copyright. The following being my conjecture (IANAL), there is fair use an…
Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?
#99Earlier quoted context omitted.
Lets turn this around the other way. I create a omniscient copyright detection bot and face it at everything you create 24 hours a day 7 days a week. You go home and sing happy birthday to your kid. The bot gives you a non-monetary warning for using a copyrighted work without permission. No big deal, but it is on your permanent record. It had been a stressful day so you take up your evening hobby of painting. You lik…
The first two situations you mention almost certainly aren’t copyright violations. The third is at least a solid “maybe.”
Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?
#100chatgpt actually has some ideas about this question: How could the people who generate used in an ai language model be paid for their work? answer: There are several ways in which the people who generate content for an AI language model could be paid for their work: Royalty-based payment: Content creators could receive a percentage of the revenue generated from the use of their content in the AI language model. Token…
[0] The answers focus on the technicalities of how the payments could be arranged, but the much bigger problem is that it's not clear who the payments should be going to (there's no immediately obvious or unique way of attributing a given output to specific training inputs, that would require a separate model with a lot of room for judgement/modelling decisions; or a new type of LLM that has that feature baked in).