Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

51–60 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#51
The importance of source citation in ChatGPT's responses is a topic of debate, particularly as the platform shifts towards a paid model. While ChatGPT is designed to deliver information in a conversational and user-friendly way, it is important to consider the potential legal implications of using unverified or uncited information. In sensitive or controversial cases, it is advisable to properly cite sources to ensure accuracy and avoid any potential issues of intellectual property infringement.

On the other hand, the focus on the potential of ChatGPT's natural language processing capabilities highlights the significance of learning and using LLM (Language Models) in data handling. The utilization of LLM can potentially lead to a future where traditional databases become obsolete and are replaced by advanced language models. As such, the development and integration of LLM in our daily lives and processes can bring about many benefits and possibilities.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#52

I agree. ChatGPT should cite its sources.

I'm pretty sure ChatGPT doesn't know it's sources. If it generated "that cat sat on the mat", then what (even from a theoretical POV) is the source of the word "mat" ? Note that it's not pulling the whole "cat sat on the mat" sentence from anyplace - that's not how it works - it's just generating this one word at a time based on the statistics (collected over all the text it was fed) of what word is most likely to follow what came before.

So, who gets credit for the word "mat" being generated in that context ? I guess any texts talking about cats and mats in close proximity may deserve some of the "credit", but it goes way deeper than that since why did ChatGPT choose to output such a trite sentence (albeit while only selecting one word at a time), rather that something else about cats or perhaps a more interesting thing that cats often sit in/on ...

People seem to assume that ChatGPT is pulling entire "facts" from various sources, but that's just not how it works - it's just feeding all the texts into a giant meat grinder of word statistics. It knows about words, not facts.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#53

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

> Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income?

I don’t do it on an industrial scale.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#54
post #30

Never realized how little data it was fed with. 570GB can fit on my laptop.

1GB file would contain roughly 166,000,000 words. This includes the space between words, so the average word is 5 characters.

A typical single-spaced page is 500 words long

That’s 179,280,000 full pages of text.

I wonder if they excluded any duplicated text.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#55
post #39

Earlier quoted context omitted.

No, it isn't that simple. The scale and totality of the scraping is out of reach for a human. If you previously interacted with people on this issue, you must know that. It is fair for a single human to breathe, but not for a machine to use all oxygen on this planet at once, killing everyone else in the process.

It is in fact that simple. There are dozens, hundreds, perhaps thousands of legitimate, genuine, serious, real reasons to be concerned and "want something to be done". This isn't it. "Learning is unfair" is not an argument you want to win.

Love how you’ve conveniently ignored

> The scale and totality of the scraping is out of reach for a human.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#57

100% agree - the likes of ChatGPT are straight up generating revenue based on adding value to stolen work.

Lets turn this around the other way.

I create a omniscient copyright detection bot and face it at everything you create 24 hours a day 7 days a week.

You go home and sing happy birthday to your kid. The bot gives you a non-monetary warning for using a copyrighted work without permission. No big deal, but it is on your permanent record.

It had been a stressful day so you take up your evening hobby of painting. You like nature scenes and trees and 30 minutes in you receive a violation, evidently Bob Ross has already done this and his surviving estate is now asking you to destroy the picture.

The next day you go to into your job at the corporate bureaucracy slinging lines of javascript. It's been a productive day so far and you have a few hundred new lines of code written and then the bots going off and HR and legal are ringing the phone within seconds. Turns out some comment you'd saw on Stack Overflow years ago was imprinted in your memory well enough you committed a copyright violation. Looks like you'll be losing your job.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#58
post #26
post #18

The data it was scraped from was then put into vector maps and usedd to create a model which is used from zero to create unique sentences that summarize what the model relates to. The text results coming out are neither copyright infringement nor plagiarism.

You're saying plagiarism isn't if one mostly swaps a couple of things in the expression of the content?

the definition of plagiarism is "the practice of taking someone else's work or ideas and passing them off as one's own"...chatGPT might infer from thousands or millions of different possible works or ideas before creating its own sentences so I don't think that meets the definition of plagiarism

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#59

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

Do you think just maybe there is a diffence here because humans need money to survive, and maybe we should have compassion for humans who could hypothetically starve or freeze or suicide or whatever because they have no money? Or is it just silly to care about people like that?

This is a response to an argument the GP didn't make. One can still have grave concerns about generative AI's potential impact on human society while accepting there is nothing fundamentally unfair about how it scrapes publicly accessible data.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#60

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

Do you think just maybe there is a diffence here because humans need money to survive, and maybe we should have compassion for humans who could hypothetically starve or freeze or suicide or whatever because they have no money? Or is it just silly to care about people like that?

That's got nothing to do with whether or not it is "fair" for a learning system to produce content after it has learned.

That is, instead, one of the larger and vastly more important sociocultural issues that actually warrants attention, but never receives it in sufficient degree to address the problem, because, for example, we're arguing whether automated learning is "fair".

Post reply on HN