Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

191–194 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#191

Earlier quoted context omitted.

The first two situations you mention almost certainly aren’t copyright violations. The third is at least a solid “maybe.”

The first situation isn't copyright violation because some monied entity went out and litigated against Warner/Chappell music. That's the problem with copyright - until you've litigated, which is expensive, you just can't tell what's in and what's out of copyright. You wrote "almost certainly" because of that.

> You wrote "almost certainly" because of that.

No, I didn’t.

I was hedging against the too-clever HN commenter coming back and saying, “The robot knew you were going to sell the painting” or “You sang the game on the Jumbotron at Yankee Stadium.”

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#192
post #185

Earlier quoted context omitted.

Why would someone only ask an LLM questions when they were in the market to buy a book? Most people I know don't buy books in order to look up the answer to a question, sure some people buy reference books and use them but that's not really what we think of when talking about authors and books. If I'm in the market for a book, I'm looking to read a book, not query something or someone for answers. I think your exampl…

> 4) They ask the LLM for a list of recent books that go in depth on the topic or are in the genre etc. 5) Your name comes up in the list My belief is that ChatGPT is actually not quite capable of that, after seeing examples of how it manufactures non-existing references. Besides, if it were capable of that, why would it not show your name as part of the answer already now? The cynic in me thinks it’s not capable of…

It's very much in their interest, if the information their models provide is impossible to verify then it severely limits its uses. You essentially can't use it as a source for anything that requires any type of citation or reliability. That's a huge handicap for selling it to businesses and researchers. The general problem of determining what training data was used to produce an output is an open problem in ML and one that is being very actively worked on since it would greatly further the field.

You believe correctly that ChatGPT is not capable of showing sources, it's currently impossible to do but we were discussing Tomorrow so I included it as a possibility. You could potentially hack it in now using traditional search or nearest neighbours but it wouldn't be 100% accurate, probably not even 50%, it would just show a bag of similar texts so not really worth doing.

I'd still be in the market for a book even if we had a perfect LLM that could answer every question I had with impeccable accuracy. I read books because I want to find out about things I don't know that I don't know. It's pretty hard to find those things if you just do question response. It's like a graph, if you start at one node it may take you a very long time to traverse the graph to another node but if you have some outside source that gives you the address of a new node you can just jump straight to it.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#193

Earlier quoted context omitted.

The first situation isn't copyright violation because some monied entity went out and litigated against Warner/Chappell music. That's the problem with copyright - until you've litigated, which is expensive, you just can't tell what's in and what's out of copyright. You wrote "almost certainly" because of that.

> You wrote "almost certainly" because of that. No, I didn’t. I was hedging against the too-clever HN commenter coming back and saying, “The robot knew you were going to sell the painting” or “You sang the game on the Jumbotron at Yankee Stadium.”

Because of the litigation, even singing "Happy birthday to you" on the Jumbotron is not infringement. But only because of the litigation.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#194
post #11

Do you want companies to do this in private for private gain and not share it with you? Because making it illegal will just make it happen in greater secrecy.

This is essentially a defeatist argument flirting with supporting extortion, it seems to me. If you think that chatgpt is doing something wrong, this is arguing that you should allow the wrong to exist because there's nothing you can do about it. In other areas of society where a bad thing cannot be stopped, we still use legislation to reduce the amount of it and mitigate some of the harm.

I don't think they're doing anything wrong. Indeed, I think they're performing a public service that none of the others seemed positioned to do. The default is to keep advances private.
Post reply on HN