Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

171–180 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#171
post #165

Earlier quoted context omitted.

> You can't strap a source-crediting mechanism on top of a transformers-based model after the fact. I've read that ChatGPT is not connected to the net, but if it was: Couldn't you have it do a google search (or better yet corpus search) for the string it generated and then return the most significant matches (significance by string matching, not google rank)? It would be really crude, but wouldn't this just be a hand…

Why couldn't you, as a human do that to verify it? The other day I had GPT write a rap battle between Burger King and Ronald McDonald. One of the stanzas came back: Burger King: Your burgers are plain, your buns a bore. Your clown's been around since '63, I'm sure my flame-grilled taste will leave you impressed My burgers are fresh, my fries are the best It turns out that yes, Ronald McDonald was first introduced in…

Because I want to credit not verify? Because I want to trace the flow of information?

It isn't just monetary compensation that's important here.

I come at this from the point of view of a scientist who is expected to reference ideas. Not necessarily back to their original source, but at least back to a source that can theoretically point back to another link in the chain.

Sure, I can manually search for a reference based on what ChatGPT gave me. Or someone could spend a few minutes adding a few lines of code to ChatGPT to save millions of people some minutes of time.

-----

What would be awesome is an LLM that you can feed data to, and it can then write a paper based solely on the data you feed it.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#172

It is not breaking the ad-based model—it’s breaking open information sharing culture as we know it. Yesterday: 1) You do research, you publish a book, you write some posts. 2) People discover your work and you personally, they visit your posts and subscribe to you. 3) You have an opportunity to upsell your book and make money on ads to sustain your future work; more importantly, you get to see traffic stats and see w…

That’s my main fear. Not the fairness / unfairness but that people might be less willing to share info and a lot becomes inaccessible / secret.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#173
post #165

Earlier quoted context omitted.

Why couldn't you, as a human do that to verify it? The other day I had GPT write a rap battle between Burger King and Ronald McDonald. One of the stanzas came back: Burger King: Your burgers are plain, your buns a bore. Your clown's been around since '63, I'm sure my flame-grilled taste will leave you impressed My burgers are fresh, my fries are the best It turns out that yes, Ronald McDonald was first introduced in…

Because I want to credit not verify? Because I want to trace the flow of information? It isn't just monetary compensation that's important here. I come at this from the point of view of a scientist who is expected to reference ideas. Not necessarily back to their original source, but at least back to a source that can theoretically point back to another link in the chain. Sure, I can manually search for a reference b…

I've still got the question - who should I credit with the bit that Ronald has been around since 1963?

I had it write a poem the other day in the style of Roses are read about coffee and bacon.

   Roses are red
   Bacon is greasy
   My coffee is hot
   Together they please me
If this is something that someone considers to be a derivative work of other things... who do I credit?

    Identify a word that have different meanings to two different professions at the same time and the professions that use them.  Give the definition of the word for each profession. Write a joke using this word.
to which I got back:

    The word is "band." 

    Definition for a Musician: A group of musicians who play music together.
    Definition for an Astronomer: A dark region in the sky with less stars.

    Joke: What did the astronomer say when the musician asked him to join his band? "I'm sorry, I don't do solos in the dark!"
How do you credit that?

---

> What would be awesome is an LLM that you can feed data to, and it can then write a paper based solely on the data you feed it.

https://platform.openai.com/docs/guides/fine-tuning

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#174
I kind of agree with you, but I think that's only because we've all been saturated with the idea of everlasting ownership of ideas.

Clearly, ownership of ideas runs out, because we all use linked lists or binary trees, or paper, or turbines or the list goes on. We don't pay money to the inventors of linked lists, or the heirs or successors-in-interest to the inventor of paper. Why not? When does ownership of an idea expire? Why do we unconsciously accept copyright or patent limits of today?

There's also an issue with simultaneous invention, but that's out of scope here. Clearly ChatGPT is just regurgitating or otherwise emitting previously-ingested material.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#175

It is not breaking the ad-based model—it’s breaking open information sharing culture as we know it. Yesterday: 1) You do research, you publish a book, you write some posts. 2) People discover your work and you personally, they visit your posts and subscribe to you. 3) You have an opportunity to upsell your book and make money on ads to sustain your future work; more importantly, you get to see traffic stats and see w…

Exactly. As a professional artist, I am expected to have a public online portfolio and publicly available imagery of shows and exhibits. Saying that I'm forfeiting my stake in my art because I'm showing it publicly is a really great way to kill art and culture. AI is not learning to make, draw, use mediums in a skilled manner. AI is scraping my public images and plottlining them with the input of humans to label them, tag them and apply stylistic qualities to them.Just because there are massive amounts of data to dilute influence doesn't change that the computer is still simply doing what a human is telling it to do with imagery created by humans. If you took away the human input, labeling and tagging you will find that the computer has not learned anything. I can look at 'AI' art and pick out artists from the collated imagery. Unlike 'AI'I can't spit out the imagery by photocopying/plottlining/tracing it. I have to learn the skills of each artist involved to recreate what I see. Motor skills require practice and effort. 'AI' is not learning motor skills, which is the basis of the creation of art. It is mapping and applying statistical algorithms to amalgamate data from preexisting sources for those who want 'Art' without the effort of time or skill to produce it. At this very moment 'AI' art is being used to sell merchandise with zero credit or monies going to the people who used their human motor skills to create the backbone of this art. Sadly,this only agravates the ways copyright already restricts human art.Imagine if we lived in a world where people valued artists with respect for thier craft? I once had someone ask me how long it took me to draw a charcoal drawing. The short answer is half an hour. The long answer is that I was doing daily scketching practice and investing many hours a week doing charcoal excercises. I am currently out of practice with charcoal and as it is a medium with no erasing or margin of error, I doubt I could recreate my drawing myself without 'getting my hand back in'. It is obvious to me that this 'AI' tool is being used by humans, with the industry of humans, to exploit humans for the gratification of end user humans. I suppose humans could stop making art to feed the monster...

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#176
post #173

Earlier quoted context omitted.

Because I want to credit not verify? Because I want to trace the flow of information? It isn't just monetary compensation that's important here. I come at this from the point of view of a scientist who is expected to reference ideas. Not necessarily back to their original source, but at least back to a source that can theoretically point back to another link in the chain. Sure, I can manually search for a reference b…

I've still got the question - who should I credit with the bit that Ronald has been around since 1963? I had it write a poem the other day in the style of Roses are read about coffee and bacon. Roses are red Bacon is greasy My coffee is hot Together they please me If this is something that someone considers to be a derivative work of other things... who do I credit? Identify a word that have different meanings to two…

> If this is something that someone considers to be a derivative work of other things... who do I credit?

Based on a quick search the best credits would be ChatGPT as the arranger, and "Roud Folk Song Index number 19798" as the inspiration.

> "Joke: What did the astronomer say when the musician asked him to join his band? "I'm sorry, I don't do solos in the dark!""

> "How do you credit that?"

That you credit to ChatGPT. It's not referencing facts or discoveries, so credit isn't as important as it is for articles. If you want to credit an inspiration then I'm sure there's an index of joke forms out there that has an appropriate number to cite.

I can't actually find a definition for band in astronomy that is "a dark region in the sky with less stars." So it seems to be a pretty poor joke.

> https://platform.openai.com/docs/guides/fine-tuning

This does it solely based on the data you feed into it? And by data I mean scientific data that you discovered, and want formatted into a particular research article style.

Edit to add: Possible sources for the line "together they please me":

1) https://www.google.com/books/edition/Poetical_Works_of_Louis...

2) https://www.google.com/books/edition/Florio_s_First_fruites/...

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#177
post #57

Earlier quoted context omitted.

Lets turn this around the other way. I create a omniscient copyright detection bot and face it at everything you create 24 hours a day 7 days a week. You go home and sing happy birthday to your kid. The bot gives you a non-monetary warning for using a copyrighted work without permission. No big deal, but it is on your permanent record. It had been a stressful day so you take up your evening hobby of painting. You lik…

The first two situations you mention almost certainly aren’t copyright violations. The third is at least a solid “maybe.”

The first situation isn't copyright violation because some monied entity went out and litigated against Warner/Chappell music. That's the problem with copyright - until you've litigated, which is expensive, you just can't tell what's in and what's out of copyright. You wrote "almost certainly" because of that.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#178
post #173

Earlier quoted context omitted.

I've still got the question - who should I credit with the bit that Ronald has been around since 1963? I had it write a poem the other day in the style of Roses are read about coffee and bacon. Roses are red Bacon is greasy My coffee is hot Together they please me If this is something that someone considers to be a derivative work of other things... who do I credit? Identify a word that have different meanings to two…

> If this is something that someone considers to be a derivative work of other things... who do I credit? Based on a quick search the best credits would be ChatGPT as the arranger, and "Roud Folk Song Index number 19798" as the inspiration. > "Joke: What did the astronomer say when the musician asked him to join his band? "I'm sorry, I don't do solos in the dark!"" > "How do you credit that?" That you credit to ChatG…

Why did you pick that index rather than some other source material? Roses are red dates back to 1784 (year not index number) as a nursery rhyme. Does it need to be credited or is it in the public consciousness to the point where one can create a poem based on it without knowing its original source?

    Write a haiku about bacon and coffee.  Identify the syllable count for each word and line used in the haiku.
    Example:
    Bacon (2) sizzles (2)
    Aroma (3) of (1) coffee (2) too (1)
    Mouthwatering (4) bliss (1)

    Smoky (2) bacon (2)
    Brewing (3) coffee (2) aroma (3)
    Makes (1) mornings (2) bright (2)
The second poem is from GPT. Do we need to credit the dictionary where it got the syllable count for each word? Or where it got that coffee (rather than bacon) is brewed? Or that bacon and coffee are things more often consumed in the morning?

    Identify four foods or beverages that are frequently consumed in the morning and how each is prepared for breakfast.

    1. Coffee: prepared by brewing hot water over ground coffee beans.
    2. Cereal: prepared by pouring cereal into a bowl and adding milk.
    3. Toast: prepared by toasting bread and adding butter and/or jelly.
    4. Eggs: prepared by scrambling, frying, poaching, or boiling them.
There is a difference between "identifying a source where this information can be found" and "this is the (copyrighted) source of the data that GPT used to draw upon to come up with the statement."

The first is an exercise for the reader (and much better done and evaluated by the reader). The second is what people are concerned about.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#179

Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income? People get internet hostile at me for this question, but it really is that simple. They've automated you, and it's definitely going to be a problem, but if it's acceptable for your brain to do the same thing, you're going to have to find a different ang…

>Is it unfair for you to create content/products/etc after you have read and learned from various sources on the internet, potentially depriving them of clicks/income?

The difference is in scale.

A human video game designer can consume other' people's art, then sell their labor to a video game developer. The amount of value captured by the video game designer rounds down to zero in terms of percentage of economic value created by 'video game art'.

OpenAI can consume all of the video game artists, ever, create an art design product and capture a significant percentage of the economic productivity of video game art.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#180
post #169

Earlier quoted context omitted.

People are already being paid to curate data for models, I’m mostly suggesting that that might become a major revenue source, and that ads might be less relevant in a world where people don’t need to sift through the trash heap to get info (and that’s a good thing overall!)

Wouldn't an AI-driven search engine be even better than a language model for that purpose though? It could even snippet highlight the most relevant parts of various web pages to save on the sifting.

Maybe, and arguably, that's what Google has been doing. But one thing I really like about the idea of using a model directly is that there's one interface to learn, whereas with web search, I'm constantly adapting to a grab bag of page types and UX conventions.
Post reply on HN