Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

181–190 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#181
post #131

Earlier quoted context omitted.

Google makes money when you click links and visit webpages. Instant info features are useful but do not directly bring Google money.

That's my point? If the product is scraping the data and presenting it on their website like ChatGPT and Google, then that's effectively the same as taking away the ad revenue from those websites because they aren't getting the impressions.

You're confused. Where there is an ad impression (a user clicking to go to a webpage from a Google search result), that webpage pays Google for bringing them traffic. If the user never clicks the ad because Google directly presented the info natively, Google doesn't make any money.

I can only make the guess that Google offers this to remain competitive against Bing; it both reduces their income and increases their tech stack.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#182
post #178

Earlier quoted context omitted.

> If this is something that someone considers to be a derivative work of other things... who do I credit? Based on a quick search the best credits would be ChatGPT as the arranger, and "Roud Folk Song Index number 19798" as the inspiration. > "Joke: What did the astronomer say when the musician asked him to join his band? "I'm sorry, I don't do solos in the dark!"" > "How do you credit that?" That you credit to ChatG…

Why did you pick that index rather than some other source material? Roses are red dates back to 1784 (year not index number) as a nursery rhyme. Does it need to be credited or is it in the public consciousness to the point where one can create a poem based on it without knowing its original source? Write a haiku about bacon and coffee. Identify the syllable count for each word and line used in the haiku. Example: Bac…

I'm concerned about both. I'm a "people".

> Why did you pick that index rather than some other source material?

I told you why references were important in scientific documents already.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#183
post #178

Earlier quoted context omitted.

Why did you pick that index rather than some other source material? Roses are red dates back to 1784 (year not index number) as a nursery rhyme. Does it need to be credited or is it in the public consciousness to the point where one can create a poem based on it without knowing its original source? Write a haiku about bacon and coffee. Identify the syllable count for each word and line used in the haiku. Example: Bac…

I'm concerned about both. I'm a "people". > Why did you pick that index rather than some other source material? I told you why references were important in scientific documents already.

Scientific documents - certainly. If you are writing a research paper or encyclopedia, I expect it to be well cited.

If you are writing something that is synthesizing knowledge (not just reporting the facts), the "where are all the places were that knowledge came from" is an impossible task for human or machine.

If I ask GPT to create a poem in the style of Roses are Red about coffee and bacon - why should that request need to be citied to the same degree of scrutiny as an encyclopedia or research paper?

If, on the other hand, you're trying to use GPT to write such a paper... I would hold that you're doing it wrong. It doesn't do that well. The model is "about" transforming language. To do so, it has a fair bit of 'knowledge' that it contains to be able to do that accurately. OpenAI makes no claims about the accuracy of the content that GPT produces (its improved, it can more accurately answer data - but if you want to know the answer it is no better than your next door neighbor who has read a lot).

If you are claiming that the example of Bacon is Greasy poem that GPT wrote is infringing any more than a child's "roses are red, my cat is orange, his eyes are green, nothing rhymes with orange" then I believe you will face an uphill battle.

To say that there is plagiarism and infringement going on - it needs examples rather than a "I think it works this way and is just regurgitating material it was fed from elsewhere."

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#184
post #131

Earlier quoted context omitted.

That's my point? If the product is scraping the data and presenting it on their website like ChatGPT and Google, then that's effectively the same as taking away the ad revenue from those websites because they aren't getting the impressions.

You're confused. Where there is an ad impression (a user clicking to go to a webpage from a Google search result), that webpage pays Google for bringing them traffic. If the user never clicks the ad because Google directly presented the info natively, Google doesn't make any money. I can only make the guess that Google offers this to remain competitive against Bing; it both reduces their income and increases their te…

It seems you are the one confused. If you present any ads on your website (Google or not) then you have ad revenue. The less traffic that comes to your website, the more that impacts your traffic thus directly affecting your ad revenue.

The original topic is about taking away money from the sources. Google taking your data and presenting it is taking away traffic because there is less traffic to the website of the data source due to less ad revenue.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#185

It is not breaking the ad-based model—it’s breaking open information sharing culture as we know it. Yesterday: 1) You do research, you publish a book, you write some posts. 2) People discover your work and you personally, they visit your posts and subscribe to you. 3) You have an opportunity to upsell your book and make money on ads to sustain your future work; more importantly, you get to see traffic stats and see w…

Why would someone only ask an LLM questions when they were in the market to buy a book? Most people I know don't buy books in order to look up the answer to a question, sure some people buy reference books and use them but that's not really what we think of when talking about authors and books. If I'm in the market for a book, I'm looking to read a book, not query something or someone for answers. I think your example should go like this:

Tomorrow: 1) you do research, write posts, publish a book, 2) it is all consumed by a for-profit operated LLM. 3) People ask LLM to get answers to some related question or interest 4) They ask the LLM for a list of recent books that go in depth on the topic or are in the genre etc. 5) Your name comes up in the list 6) Goto step 2 from Yesterday

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#186
post #144
post #63

Earlier quoted context omitted.

> The question is if this data is legal to scrape ...it is? I didn't see that question raised in OP's text at all. What do legacy human legalities have to do with how AI will behave? > Because it's false equivalence? ChatGPT isn't a human being. Is this important? What is so special about human learning that it puts it in a morally distinct category from the learning that our successors will do? It sounds like OP is…

>Is this important? Well yes, it's the whole crux of the matter. Laws govern human behaviour. As of 2023, only living beings have agency. If I shoot someone with a gun, the criminal is me and not the gun. Being a deterministic piece of silicon, a computer is perfectly equivalent. Sure, it is important to start a discussion of potential nonhuman sentience in the future, but these AI models are not unlike any previous…

> these AI models are not unlike any previous software in legal issues

Agreed. However, the previous 'legal issues' related to software and the emergence of the internet are also difficult to take seriously when considered on anything but extremely short time scales.

Every time we swirl around this topic, we arrive at the same stumbles which the legacy legal system refuses to address:

* If something happening on the internet is illegal, _where_ is it illegal? Different jurisdictions recognize different jurisdictional notions - they can't even agree on whose laws apply where. If you declare something to be illegal in your house, does that give it the force of law on the internet? Of course not. Yet, the internet doesn't recognize the US state any more than it does your household. It seamlessly routes around the "laws" of both.

* The "laws" that the internet is bound to follow are the fundamental forces of physics. There is no - and can be no - formal in-band way for software to be bound to the laws of men, because signals do not obey borders. The only way to enforce these "laws" are out-of-band violence.

* States continuously, and without exception, find themselves at a disadvantage when they make the futile effort to stem the evolution of the internet. For example, only 30 years ago (a tiny spec in evolutionary time scales), the US state gave non-trivial consideration to banning HTTPS.

I understand that people sometimes follow laws. But they also often don't. The internet has already formed robust immunity against human laws.

Whatever human laws are, they are not the crux of anything related to evolution of software. They are already routinely cast aside when necessary, and are very clearly headed for total irrelevance.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#188
post #185

It is not breaking the ad-based model—it’s breaking open information sharing culture as we know it. Yesterday: 1) You do research, you publish a book, you write some posts. 2) People discover your work and you personally, they visit your posts and subscribe to you. 3) You have an opportunity to upsell your book and make money on ads to sustain your future work; more importantly, you get to see traffic stats and see w…

Why would someone only ask an LLM questions when they were in the market to buy a book? Most people I know don't buy books in order to look up the answer to a question, sure some people buy reference books and use them but that's not really what we think of when talking about authors and books. If I'm in the market for a book, I'm looking to read a book, not query something or someone for answers. I think your exampl…

> 4) They ask the LLM for a list of recent books that go in depth on the topic or are in the genre etc. 5) Your name comes up in the list

My belief is that ChatGPT is actually not quite capable of that, after seeing examples of how it manufactures non-existing references. Besides, if it were capable of that, why would it not show your name as part of the answer already now?

The cynic in me thinks it’s not capable of that primarily because it is not a priority for OpenAI and training data strips attribution, with an explicit purpose: if the public knows that ChatGPT can trace back the source, OpenAI would be on the hook for paying all the countless non-consensual content providers on which work it makes money.

We should treat OpenAI as we treat Google and Microsoft. It has great talent and charismatic people working for it, but ultimately it’s a for-profit tech company and the name they chose ought to make us all the more suspicious (akin to Google’s “don’t be evil”).

> Why would someone only ask an LLM questions when they were in the market to buy a book?

Why would you be in a market for a book when you can learn the same and more by asking an LLM that already consumed said book? And therefore why would the author spend effort writing and publishing a book knowing it’d sell exactly one copy (to LLM operator)?

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#189

It is not breaking the ad-based model—it’s breaking open information sharing culture as we know it. Yesterday: 1) You do research, you publish a book, you write some posts. 2) People discover your work and you personally, they visit your posts and subscribe to you. 3) You have an opportunity to upsell your book and make money on ads to sustain your future work; more importantly, you get to see traffic stats and see w…

That’s my main fear. Not the fairness / unfairness but that people might be less willing to share info and a lot becomes inaccessible / secret.

I am also anxious about the web becoming fragmented and secretive. If one must gain access to the right circles to start learning, it hinders learning in general, and for myself and many people I know would basically mean we wouldn’t be doing what we’re doing if it were the case when we were younger.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#190

Earlier quoted context omitted.

That’s my main fear. Not the fairness / unfairness but that people might be less willing to share info and a lot becomes inaccessible / secret.

I am also anxious about the web becoming fragmented and secretive. If one must gain access to the right circles to start learning, it hinders learning in general, and for myself and many people I know would basically mean we wouldn’t be doing what we’re doing if it were the case when we were younger.

Exactly. We’re in ChatGPT honeymoon but incentives to share info moving forward unclear. Could see big model owners paying for exclusive access to content / data hindering the free distribution of information and becoming like the old publishers and gatekeepers.
Post reply on HN