Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

11–20 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#13
Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income?

People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different angle to attack it than "fairness".

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#14

Well you could say same thing about the answers that Google displays on it's pages instead of search results! If you don't want these crawlers to index your content I am pretty sure you can disable via robots.txt just like Google.

The training model data sets have inconsistent respect for robots.txt. Also, I believe most of these models are not continuously crawling websites to update their data like a search engine does. That means if you're crawled once, you may not be crawled again and you'll still be in the datasets.

I'd also argue that Google directing traffic to your website is a good alignment of incentives. ChatGPT spitting out answers derived from your work with nothing given back to you in return is not.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#15
You can also invert this and say that without a system like ChatGPT it is physically impossible for most people to find or use those 570GB of data. A search engine can only get you so far and over time they are becoming less useful as the net floods with junk content. If you don't even know what terms to search for then ChatGPT wins out since you can start with a very simple question and then interrogate it further on details it produces. The best way to think about it is as a better search engine, a fully interactive one that also has some degree of its own agency when it comes to synthesizing data. It could be better, it would be nice to have the option to show sources for the output so that you can verify the facts or do your own research.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#16

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

You usually expect people to cite sources. Granted, that very often doesn't happen, and the amount of citing expected depends on the context. But ChatGPT just doesn't cite sources at all. I think there's a case to be made that they should.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#18
The data it was scraped from was then put into vector maps and usedd to create a model which is used from zero to create unique sentences that summarize what the model relates to. The text results coming out are neither copyright infringement nor plagiarism.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#20
post #2

Is it unfair that to present coworkers thoughts you summarized or derived after reading ad-supported content?

Is that something you think is a good analogy for ChatGPT use of it's data sources?

To me it looks more like memorizing enough of other employees' project contributions to try passing it all off as your own achievements in performance review.

Post reply on HN