Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

41–50 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#41
The dirty secret of how so many social media giants got their initial traction in the early growth stage they scraped content. LinkedIn is one I have personal knowledge of. Facebook another. How do you think they got a critical mass of users? Scraping and fake engagement. Back in the 00's when they were startups operating in little offices in the SF Bay, they had teams of people running Beautiful Soup and were building bots to build profiles and stuff.

I'm actually not really sure I have an opinion on the ethics of it. Same argument as Adblock. You don't get to control how people consume your content if you put it out in the world for free. That goes for profiles, or articles, reddit posts, StackOverflow, etc. The only thing that's ironic is that large tech companies throw a fit whenever you want to turn the tables and scrape them.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#42
post #21

Earlier quoted context omitted.

The idea that a robots.txt will save you is laughable.

Agreed. At best, you can disallow: / and hope they're polite enough to listen. I can't seem to find anything on OpenAI's crawler agent, so I'm skeptical they're considering robots.txt at all.

Even if they abide, this is capitalism. Somebody who wants an edge won't. Or OpenAI or Google will get desperate and stop abiding.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#44

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

> Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income?

Because it's false equivalence? ChatGPT isn't a human being. It's a product that is built upon data from other sources.

The question is if this data is legal to scrape, which it is: Web scraping is legal, US appeals court reaffirms [https://news.ycombinator.com/item?id=31075396].

As long as the content is not copyrighted and it's not regurgitating the exact same content, then it should be okay.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#45

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

At a human level it falls below the noise floor. It's a fact of life that humans will learn and build from experience. The difference is scale. At scale it becomes a problem. Edit: I don't know how to satisfy all parties. This shakes the foundation of copyright. Perhaps we are all finding out how valuable good information truly is and especially in aggregate. We have created proto-gods.

At scale, it becomes a wonderful tool. Are the people in this thread so threatened or so invested in the current business models of the internet that you can’t see how amazing this sort of thing could be for our abilities as a species? Not just in its current iteration, but it will get better and better.

This could be an excellent brain augmentation, trying to hamper it because we want to force people to drag themselves through underlying sources so those sources can try to steal their attention with ads for revenue is asinine.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#46

The dirty secret of how so many social media giants got their initial traction in the early growth stage they scraped content. LinkedIn is one I have personal knowledge of. Facebook another. How do you think they got a critical mass of users? Scraping and fake engagement. Back in the 00's when they were startups operating in little offices in the SF Bay, they had teams of people running Beautiful Soup and were buildi…

Didn’t LinkedIn use people’s phone books

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#47

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

Do you think just maybe there is a diffence here because humans need money to survive, and maybe we should have compassion for humans who could hypothetically starve or freeze or suicide or whatever because they have no money? Or is it just silly to care about people like that?

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#48
post #45

Earlier quoted context omitted.

At a human level it falls below the noise floor. It's a fact of life that humans will learn and build from experience. The difference is scale. At scale it becomes a problem. Edit: I don't know how to satisfy all parties. This shakes the foundation of copyright. Perhaps we are all finding out how valuable good information truly is and especially in aggregate. We have created proto-gods.

At scale, it becomes a wonderful tool. Are the people in this thread so threatened or so invested in the current business models of the internet that you can’t see how amazing this sort of thing could be for our abilities as a species? Not just in its current iteration, but it will get better and better. This could be an excellent brain augmentation, trying to hamper it because we want to force people to drag themsel…

It is a wonderful tool but I still feel that the creators of the training data are getting shafted. I'm both amazed and horrified at our creation and what it portends.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#49
post #39

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

No, it isn't that simple. The scale and totality of the scraping is out of reach for a human. If you previously interacted with people on this issue, you must know that. It is fair for a single human to breathe, but not for a machine to use all oxygen on this planet at once, killing everyone else in the process.

It is in fact that simple. There are dozens, hundreds, perhaps thousands of legitimate, genuine, serious, real reasons to be concerned and "want something to be done". This isn't it.

"Learning is unfair" is not an argument you want to win.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#50
post #6

Earlier quoted context omitted.

It's the lack of attribution that really hurts, though I think its fairly shady of google to steal the ad revenue from smaller sites.

You can ask ChatGPT to cite its sources.

You can ask it to, but it will just make up sources. The connection between its knowledge and the original sources is not represented in the model.

(This is an active area of research, though, and version of GPT that could cite its sources is something people widely agree would be valuable.)

Post reply on HN