Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

161–170 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#162
post #137

Earlier quoted context omitted.

It is a wonderful tool but I still feel that the creators of the training data are getting shafted. I'm both amazed and horrified at our creation and what it portends.

Yeah, there will probably have to be some adjustment. In the future, maybe an ML agent will hire people to go find answers for it about questions it has, using us as researchers/mechanical Turks :-) Quality matters more than quantity for something that’s trying to understand the world well and not just building a statistical language model, I imagine that it will be worth it to pay for quality when training heavily u…

The internet has always been a trash heap. We've just been creating new heaps with parts of the old heaps every few years or so. Sure, it's nice to imagine a future in which this isn't the case, but your imagination is not going to be the future reality.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#163
post #39

Earlier quoted context omitted.

No, it isn't that simple. The scale and totality of the scraping is out of reach for a human. If you previously interacted with people on this issue, you must know that. It is fair for a single human to breathe, but not for a machine to use all oxygen on this planet at once, killing everyone else in the process.

If I woke up tomorrow and breathed all the oxygen, nobody else could breathe. But If I woke up tomorrow and read all the websites on the internet, it wouldn't stop other people from reading them too. Air is zero-sum. Knowledge is not.

> But If I woke up tomorrow and read all the websites on the internet, it wouldn't stop other people from reading them too.

If you became the first line, go-to source for the information of those websites, those websites would stop getting click-throughs. Eventually it would become less and less worthwhile (economically or emotionally) for the people keeping those sites running to keep them running. It would become more and more difficult for people to find those sites even if they are running, or even the archives of those sites.

So yes, eventually you'd stop people from reading them too.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#164
post #16

Earlier quoted context omitted.

You usually expect people to cite sources. Granted, that very often doesn't happen, and the amount of citing expected depends on the context. But ChatGPT just doesn't cite sources at all. I think there's a case to be made that they should.

Humans have a pretty good sense of when you need to cite sources, and when you don't. For example, long ago I learned from some website how to write a for-loop in python, and now I write them all the time without giving credit. I'm okay with ChatGPT writing a for-loop without citing its source. I would say most knowledge about words/grammar/laws of nature can be taken for granted without a citation, but there are som…

And yet, exactly in this example, I HATE that people don't put sources. Perhaps not for "for loops", but search anything simple in python. "Python JSON output", for example, and you will find a billion articles that describe a simple python library ... but DON'T link to python.org or the "javadoc". They're always dicussing the most blatantly obvious simple thing, never remotely complete, never link to where you can actually find more info (but jobs, courses, ads, ... those will be linked)

It's getting me to the point of refusing to use Google, or only use Google with "site:...". I mean, the site varies, but without site limits Google's becoming useless.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#165

Earlier quoted context omitted.

ChatGPT doesn't have a concept of sources. It has weights that together define a function that allow it to guess the most likely next word from the context. As a neat side effect of this contextual next-word guessing, it often can share accurate information. If ChatGPT were to be required to share its sources, they would need a completely different approach. I'm not commenting on whether or not that would be a bad th…

> You can't strap a source-crediting mechanism on top of a transformers-based model after the fact. I've read that ChatGPT is not connected to the net, but if it was: Couldn't you have it do a google search (or better yet corpus search) for the string it generated and then return the most significant matches (significance by string matching, not google rank)? It would be really crude, but wouldn't this just be a hand…

Why couldn't you, as a human do that to verify it?

The other day I had GPT write a rap battle between Burger King and Ronald McDonald. One of the stanzas came back:

    Burger King:
    Your burgers are plain, your buns a bore.
    Your clown's been around since '63,
    I'm sure my flame-grilled taste will leave you impressed
    My burgers are fresh, my fries are the best
It turns out that yes, Ronald McDonald was first introduced in 1963. https://en.wikipedia.org/wiki/File:McDonald%27s_commercial_(... (from https://en.wikipedia.org/wiki/Willard_Scott#Created_Ronald_M... )

So here's the challenge for you - who do you compensate for that line?

The complaint that people have isn't that GPT isn't citing its sources but rather that it isn't compensating the people who created the data that has that information.

... and now, if you're ever asked about historical clown trivia and pull out the "Ronald has been around since 1963", who should you give a royalty to? Me (for writing this), GPT (for making me aware of it), Wikipedia (for the source of my links in this post), the estate of Willard Scott for the Joy of Living (which Wikipedia cites), some random blog author that had some clown trivia on it that happened to have been part of the training set for GPT?

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#166
post #104

Absolutely, yes. It's incredibly unfair. But techbros here and elsewhere don't care about you or me or people in general and they'll think up an infinite amount of ridiculous false equivalencies before admitting the risks and real harms.

> but techbros here and elsewhere don't care about you or me or people in general and they'll think up an infinite amount of ridiculous false equivalencies before admitting the risks and real harms.

I came to this realisation arguing with someone in a mutual discord server, about these very topics (the negative impacts of AI). They just couldn't see it, and refused to believe it. I was constantly met with things like "Sure, we'll have to adjust but it'll come" and "Things are no worse now than when the TV and when books were invented" (completely ignoring the many of billions companies are spending to make things more addictive ot our monkey minds, which don't change). Also lots of noble "everyone can use it and it'll benefit everyone"...when really, it only benefits those who can control it. No mention of biases in training data or anything else either. They were really completely blinded to the idea that it might not be good and we should serious admit there are huge issues looming.

I also found it telling that the multiple people like that also weren't fans of in-person interaction, outside their friend group. They saw Discord interactions as just as fine as going out and having serendipitous moments in person, with other real people, and just actually living. Something else I feel technology has stolen from us with everyone always glued to their screen. It's funny how I've become something of a Luddite, proudly, and think we need less internet and more real world, cause, well, life is real world, being human is through real world interactions. And not ones mediated by your phone.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#167
post #92

Earlier quoted context omitted.

ChatGPT isn't doing the scraping, humans are. And humans are using computers to both read the article and create content or to scrape it. So not it's not a false equivalence.

There’s a reason scraping is a legally grey area. > Web scraping is legal, US appeals court reaffirms First, the case is not closed. [0] Second, to draw an analogy, you can use scraping in the same way you can use a computer: for legal purposes. That is, you cannot use scraping to violate copyright, just as you cannot use a computer to violate copyright. The following being my conjecture (IANAL), there is fair use an…

It seems the issue with scraping as it pertains to copyright issues isn't the scraping, any more than buying a book to sell off photocopies of it cheaply doesn't indicate that there is a problem with buying books. The issue is the copying, and more importantly, the distribution of those copies.

Fair use of course being the exception.

Now, as for accessing things like credentials that get left in unsecured AWS buckets is the bigger area where courts are less likely to recognize the legality of scraping. Never mind the fact that these people literally published their private data on a globally accessible platforms in a public fashion. I'm not a lawyer but I've seen reports of this leaning both directions in court, and yes, I've seen wget listed as a "hacker tool."

This is what happens when feelings matter more to the legal system than principles.

And before it's brought up, I may as well point out that no, I don't condone the actual USE of obviously private credentials found in an AWS bucket any more than I condone the use of a credit card that one may find on the sidewalk. Both are clearly in the public sphere, unprotected, but for both there is a pretty good expectation that someone put it there by accident, and that it's not YOUR credential to use.

Basically, getting back to the OP, ChatGPT hasn't done anything I've seen that'd constitute copyright infringement -- fair use seems to apply fairly well. As for the ad-supported model, adblockers did this all first. If you wanted to stop anything accessing your site that didn't view ads, there are solutions out there to achieve this. Don't be surprised when it chases away a good amount of traffic though -- you're likely serving up ad-supported content because it's not content you expected your users to pay for to begin with.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#168
post #104

Absolutely, yes. It's incredibly unfair. But techbros here and elsewhere don't care about you or me or people in general and they'll think up an infinite amount of ridiculous false equivalencies before admitting the risks and real harms.

1. get to enjoy an open network of networks

2. people share, get creative and get some sort of credit for it

3. scrap it all and feed it a large deep neural network and be a worse version of all this content but easily accessible

4. creative people don't see a reason to keep sharing what they have (no new public books, no new open source projects, ...)

5. get stuck in an AI world of recycled content

People blindly following OpenAI products have a very shortsighted vision. What they did is neither innovative, nor extraordinary, they got the data, convinced some victims into a kickstart, made sure the hardware supports the bigger deep neural network that can do the job. Check out the OpenAI alternative solutions, it's not hard.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#169
post #137

Earlier quoted context omitted.

Yeah, there will probably have to be some adjustment. In the future, maybe an ML agent will hire people to go find answers for it about questions it has, using us as researchers/mechanical Turks :-) Quality matters more than quantity for something that’s trying to understand the world well and not just building a statistical language model, I imagine that it will be worth it to pay for quality when training heavily u…

The internet has always been a trash heap. We've just been creating new heaps with parts of the old heaps every few years or so. Sure, it's nice to imagine a future in which this isn't the case, but your imagination is not going to be the future reality.

People are already being paid to curate data for models, I’m mostly suggesting that that might become a major revenue source, and that ads might be less relevant in a world where people don’t need to sift through the trash heap to get info (and that’s a good thing overall!)

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#170
post #169

Earlier quoted context omitted.

The internet has always been a trash heap. We've just been creating new heaps with parts of the old heaps every few years or so. Sure, it's nice to imagine a future in which this isn't the case, but your imagination is not going to be the future reality.

People are already being paid to curate data for models, I’m mostly suggesting that that might become a major revenue source, and that ads might be less relevant in a world where people don’t need to sift through the trash heap to get info (and that’s a good thing overall!)

Wouldn't an AI-driven search engine be even better than a language model for that purpose though? It could even snippet highlight the most relevant parts of various web pages to save on the sifting.
Post reply on HN