Live data from Hacker News

Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

news.ycombinator.com

131–140 of 194 posts

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#131
post #116

All I have to say is, as technologists, anyone who is criticizing ChatGPT and has not been criticizing Google is a hypocrite. It's well known Google tries to keep you on Google by parsing more and more information from websites and summarizing it. Ex, Wikipedia summaries, IMDB Scores, Review Stars, etc... If you have a problem with ChatGPT's "scraped data", then you have more fundamental issues with how the internet…

Google makes money when you click links and visit webpages. Instant info features are useful but do not directly bring Google money.

That's my point?

If the product is scraping the data and presenting it on their website like ChatGPT and Google, then that's effectively the same as taking away the ad revenue from those websites because they aren't getting the impressions.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#132

Earlier quoted context omitted.

The first two situations you mention almost certainly aren’t copyright violations. The third is at least a solid “maybe.”

They aren't, or they shouldn't be, but that's the point of the parent's comment. Look at the videos flagged by youtube or copyright trolls, a lot of them are not actual copyright violations, but they are flagged anyway by the algorithm and removed or demonetized. And it takes a lot of work to fight those claims.

No, that doesn’t seem to be the point of the parent’s comment. That comment is treating those activities like copyright violations when they’re not. Why would the parent imagine an omniscient copyright bot that is probably wrong?

Me singing happy birthday at home or painting a picture for myself to relax is already demonetized. There doesn’t need to be an omniscient (and wrong) copyright bot to do that.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#133
It is not breaking the ad-based model—it’s breaking open information sharing culture as we know it.

Yesterday: 1) You do research, you publish a book, you write some posts. 2) People discover your work and you personally, they visit your posts and subscribe to you. 3) You have an opportunity to upsell your book and make money on ads to sustain your future work; more importantly, you get to see traffic stats and see what is in demand, you get thank-you emails and feel valued.

Tomorrow: 1) you do research, write posts, publish a book, 2) it is all consumed by a for-profit operated LLM. 3) People ask LLM to get answers, and have no reason or even opportunity to buy your book or know you exist.

What exactly are the incentives to publish information openly in that world?

(Will they even believe you if you say you’re the one who did the niche research powering some specific ChatGPT answer, in a world everyone knows that you can just ask an LLM?)

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#134

Earlier quoted context omitted.

It is in fact that simple. There are dozens, hundreds, perhaps thousands of legitimate, genuine, serious, real reasons to be concerned and "want something to be done". This isn't it. "Learning is unfair" is not an argument you want to win.

Love how you’ve conveniently ignored > The scale and totality of the scraping is out of reach for a human.

Because it's conveniently irrelevant.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#135
post #93

When this becomes properly entrenched I fear that it may create a disincentive to create original content. If that happens we will all be poorer for it in return for amazing access to what we already have. I don't think it is a good deal.

It's my opinion that royalties are the reason we have so much horrifying junk in our culture. It has created a world where we are inundated with cultural garbage that people produced only to squeeze money out of copyrights.

I dream of royalties going away so that we only original content that was made for the love of expression, a feeling that it's important. I would be happy to have a LOT less stuff to look at if I didn't have to sift through so much garbage.

Of course, I am also in favor of UBI so that those creators can eat while they are doing it.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#136
post #30

Never realized how little data it was fed with. 570GB can fit on my laptop.

The original dataset was 45 TB.

The neural net model is condensed to 800 GB.

https://www.springboard.com/blog/data-science/machine-learni...

Note that the "compression" there also includes the "intelligence" that it presents - you might be able to get some powerful compression of English text... but you can't ask a gzip file to come up with a joke about cats and dinosaurs.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#137
post #45

Earlier quoted context omitted.

At scale, it becomes a wonderful tool. Are the people in this thread so threatened or so invested in the current business models of the internet that you can’t see how amazing this sort of thing could be for our abilities as a species? Not just in its current iteration, but it will get better and better. This could be an excellent brain augmentation, trying to hamper it because we want to force people to drag themsel…

It is a wonderful tool but I still feel that the creators of the training data are getting shafted. I'm both amazed and horrified at our creation and what it portends.

Yeah, there will probably have to be some adjustment. In the future, maybe an ML agent will hire people to go find answers for it about questions it has, using us as researchers/mechanical Turks :-) Quality matters more than quantity for something that’s trying to understand the world well and not just building a statistical language model, I imagine that it will be worth it to pay for quality when training heavily used models, to avoid using garbage info. You don’t need 30 different superficial product reviews with a bunch of SEO text if you have one that’s very thoroughly researched.

And in the meantime, with ads no longer working, maybe crypto is actually useful for something here - lightning makes very small transactions possible with basically no fees, and makes it easy to programmatically pay for things. People hate being nickled and dimed, but a professional trying to construct an ML model could reasonably budget for use fees for fast unhindered access to quality training data. An agent could even evaluate its likelihood of learning something new/accurate vs the cost proposed by the server, and choose the subsets to pull.

Just a random idea, but I hope we don’t fight tooth and nail to preserve the trash heap of the internet’s current state.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#138

Earlier quoted context omitted.

With search engines, it does feel like there was is a more clear trade of scraping access in exchange for web traffic. With ChatGPT the traffic benefit isn’t there, so it feels like it isn’t a fair trade. Google adding the context and data to their search results page also started blurring this trade making it unnecessary to click to the site the info was cleaned from.

How does someone site a source when they are using GPT to convert a box score into an entertaining paragraph about a baseball game? Or to convert a natural language command into a JSON format ready for downstream processing?

Right, it’s a gestalt from a huge set of sources. It’s not copying single text sources verbatim into your output.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#139
post #98

Earlier quoted context omitted.

Yes but that's a technical issue. I took the parent as making a philosophical point and responded in that spirit.

Wouldn’t it be nice if the people on these forums were not ignorant of both philosophy or the legal system before diving into incoherent conversations about both at the same time where the main thrust is the emotions they have about these tools?

One can dream.

Re: Ask HN: Isn't ChatGPT unfair to the sources it scraped data from?

#140

I think it should reference the sources of the information, similar to any research paper or essay.

Could you cite the need to use citations in common conversation or asking it to write jokes?

If you are using GPT as a research tool as opposed to asking your friend who is an expert int the subject, are you citing your friend when you write the paper - or are you going back and finding sources that then back your friend's point up?

Post reply on HN