Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

21–30 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#21

Well you could say same thing about the answers that Google displays on it's pages instead of search results! If you don't want these crawlers to index your content I am pretty sure you can disable via robots.txt just like Google.

The idea that a robots.txt will save you is laughable.

Agreed. At best, you can disallow: / and hope they're polite enough to listen.

I can't seem to find anything on OpenAI's crawler agent, so I'm skeptical they're considering robots.txt at all.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#22
post #16

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

You usually expect people to cite sources. Granted, that very often doesn't happen, and the amount of citing expected depends on the context. But ChatGPT just doesn't cite sources at all. I think there's a case to be made that they should.

People don’t remember the sources that formed their opinions, it’s just baked into the structure of their brain after reading, same for the model.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#23
post #6

Well you could say same thing about the answers that Google displays on it's pages instead of search results! If you don't want these crawlers to index your content I am pretty sure you can disable via robots.txt just like Google.

It's the lack of attribution that really hurts, though I think its fairly shady of google to steal the ad revenue from smaller sites.

You can ask ChatGPT to cite its sources.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#24
Big Tech companies have been scraping massive amounts of data for about two decades. Many smaller companies have tried to imitate them (remember when Big Data was the hottest thing out there? How do you think most of those startups obtained their data?) but pretty much all of them failed, mainly by running out of cash. OpenAI just happened to win the scraping lottery.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#25
I think this is a real concern, but imagine a couple other scenarios:

1. You have a widely read spouse named Joe who reads constantly. He's got a good memory, and typically if you have a question you just ask him instead of searching for it yourself. Are you depriving Joe's sources of your eyeballs?

2. Many books summarize and restate other books. If I read Cliff's Notes on a book, for example, I can learn a lot about the original book without buying it. Is this depriving the author?

3. I have a website that proxies requests to other websites and summarizes them while stripping out ads.

So which of these examples are a better metaphor for what a LLM does?

I don't know. The fact is, LLMs are a new thing in our tech and culture and they don't quite fit into any of our existing cultural intuitions or norms. Of course it's ambiguous! But it's also exciting.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#26
post #18

The data it was scraped from was then put into vector maps and usedd to create a model which is used from zero to create unique sentences that summarize what the model relates to. The text results coming out are neither copyright infringement nor plagiarism.

You're saying plagiarism isn't if one mostly swaps a couple of things in the expression of the content?
Post reply on HN