Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

151–160 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#151

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

Do you think just maybe there is a diffence here because humans need money to survive, and maybe we should have compassion for humans who could hypothetically starve or freeze or suicide or whatever because they have no money? Or is it just silly to care about people like that?

Isn't that like saying automated looms should be banned because it meant humans would lose jobs to it? Or buggy whip drivers wanting to ban cars?

https://en.wikipedia.org/wiki/Luddite

Might as well ban computers since they automated and eliminated a lot of manual jobs.

The problem of humans with no money should be solved by a societ safety net and things like UBI.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#152
post #141

Earlier quoted context omitted.

How is it not scraping? There's no other way to get all that data for training a model without scraping.

It's scraping both when humans do it and when the ChatGPT team do it, but that wasn't the point the parent made. He made a moral/philosophical point which is what i responded to.

[deleted]

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#153
post #98

Earlier quoted context omitted.

There’s a reason scraping is a legally grey area. > Web scraping is legal, US appeals court reaffirms First, the case is not closed. [0] Second, to draw an analogy, you can use scraping in the same way you can use a computer: for legal purposes. That is, you cannot use scraping to violate copyright, just as you cannot use a computer to violate copyright. The following being my conjecture (IANAL), there is fair use an…

Yes but that's a technical issue. I took the parent as making a philosophical point and responded in that spirit.

[deleted]

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#154
post #141

Earlier quoted context omitted.

How is it not scraping? There's no other way to get all that data for training a model without scraping.

It's scraping both when humans do it and when the ChatGPT team do it, but that wasn't the point the parent made. He made a moral/philosophical point which is what i responded to.

[deleted]

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#155

Anybody have an idea of what kind of hardware (and cost) you would need to train the model and to execute it ? Obviously storage is not a major factor here.

I seem to recall that the training cost for ChatGPT was in the tens of millions of dollars. Execution cost is on the order of ~$1 per interaction.

The cost per day isn't something that there are any reliable sources for.

The closest to an authoritative source on it is https://twitter.com/sama/status/1599671496636780546

> average is probably single-digits cents per chat; trying to figure out more precisely and also how we can optimize it

An attempt to work through it from related resources is https://twitter.com/tomgoldsteincs/status/160019698195510069...

In particular https://twitter.com/tomgoldsteincs/status/160019699090561433...

> So what would this cost to host? On Azure cloud, each A100 card costs about $3 an hour. That's $0.0003 per word generated.

> But it generates a lot of words! The model usually responds to my queries with ~30 words, which adds up to about 1 cent per query.

---

It is much less than $1/interaction.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#156
post #14

Earlier quoted context omitted.

The training model data sets have inconsistent respect for robots.txt. Also, I believe most of these models are not continuously crawling websites to update their data like a search engine does. That means if you're crawled once, you may not be crawled again and you'll still be in the datasets. I'd also argue that Google directing traffic to your website is a good alignment of incentives. ChatGPT spitting out answers…

I bet that fully half the time, I read the google answer, click on nothing and go on my way.

That's still better than 0%

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#157

Earlier quoted context omitted.

Do you think just maybe there is a diffence here because humans need money to survive, and maybe we should have compassion for humans who could hypothetically starve or freeze or suicide or whatever because they have no money? Or is it just silly to care about people like that?

Isn't that like saying automated looms should be banned because it meant humans would lose jobs to it? Or buggy whip drivers wanting to ban cars? https://en.wikipedia.org/wiki/Luddite Might as well ban computers since they automated and eliminated a lot of manual jobs. The problem of humans with no money should be solved by a societ safety net and things like UBI.

Well, the luddites were right in that they were fighting a good and honorable fight.

Until things are, in fact, solved by whatever idea you might have, why should we just accept each new thing that makes our human lives more intolerable? How could you expect any rational person to have that kind of blind trust in a technology, much less "progress" itself, when every single aspect of our world shows that it is who owns the technology that actually benefits from it? I think it is much more crazy just totally rolling over for each new thing that takes your job than it is to maybe fight for your food and shelter.

I think we can do better than UBI, but either way, fighting against this unfairness is fighting for the things we need to continue with some shred of humanity, insofar as this technology is and will be an agent for the consolidation of labor and profit. Its all the same fight, and the historical luddites understood this consciously or not.

Who knows, maybe the internet would have been better off if some people were brave enough to smash some of Google's servers in like 2006..

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#158
post #44

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

> Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? Because it's false equivalence? ChatGPT isn't a human being. It's a product that is built upon data from other sources. The question is if this data is legal to scrape, which it is: Web scraping is legal, US appeals court reaffirms [ https://news.ycom…

Check me on this because I'm not a software person:

When a person "scrapes" a website by clicking through the link it registers as a hit on the website and, without filters being turned on, triggers the various ad impressions and other cookies. Also if the person needs that information again odds are they'll click on a bookmark or a search link and repeat the impression process all over again.

When an AI scrapes the web it does so once, and possibly in a manner designed to not trigger any ads or cookies (unless that's the purpose of the scrape). It's more equivalent to a person hitting up the website through an archive link.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#159
post #16

Earlier quoted context omitted.

You usually expect people to cite sources. Granted, that very often doesn't happen, and the amount of citing expected depends on the context. But ChatGPT just doesn't cite sources at all. I think there's a case to be made that they should.

ChatGPT doesn't have a concept of sources. It has weights that together define a function that allow it to guess the most likely next word from the context. As a neat side effect of this contextual next-word guessing, it often can share accurate information. If ChatGPT were to be required to share its sources, they would need a completely different approach. I'm not commenting on whether or not that would be a bad th…

> You can't strap a source-crediting mechanism on top of a transformers-based model after the fact.

I've read that ChatGPT is not connected to the net, but if it was: Couldn't you have it do a google search (or better yet corpus search) for the string it generated and then return the most significant matches (significance by string matching, not google rank)? It would be really crude, but wouldn't this just be a handful of lines of code that don't interfere with the "transformers-based model" code at all?

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#160
post #144
post #63

Earlier quoted context omitted.

> The question is if this data is legal to scrape ...it is? I didn't see that question raised in OP's text at all. What do legacy human legalities have to do with how AI will behave? > Because it's false equivalence? ChatGPT isn't a human being. Is this important? What is so special about human learning that it puts it in a morally distinct category from the learning that our successors will do? It sounds like OP is…

>Is this important? Well yes, it's the whole crux of the matter. Laws govern human behaviour. As of 2023, only living beings have agency. If I shoot someone with a gun, the criminal is me and not the gun. Being a deterministic piece of silicon, a computer is perfectly equivalent. Sure, it is important to start a discussion of potential nonhuman sentience in the future, but these AI models are not unlike any previous…

> It's bizarre to me how many people are missing this.

Very much this. I am too tired right now to engage with other responders, but thank you for articulating precisely the point I want to make.

Post reply on HN