Live data from Hacker News

Big Tech's underground race to buy AI training data

reuters.com

81–90 of 152 posts

Re: Big Tech's underground race to buy AI training data

#81
post #77
post #66

Earlier quoted context omitted.

isnt that what Google did ? they scraped the internet but the public/econ advisors felt the benefits outweighed copyright violations, they were just "indexers", they weren't scraping "news" they were indexing it lol same thing with emulators and roms. somebody dumped the cartridges (copyrighted software) into ROM files to be played on emulators (copyrighted bios) but they were "archiving" and if you owned the origina…

The same benefit doesn’t exist for ChatGPT as Google because Google means people click on your site and you get ad revenue. Google even facilitates this in both directions with search ads and as an ad service you can get paid from for hosting ads. The ROM site DMCA thing was always BS lmao it’s completely legal for you to dump your own carts and use them in emulators but that freedom doesn’t extend to having a copy o…

so you think scraping copyrighted content to sell ads is okay and downloading copyrighted games for free is also okay then why is it not okay for ChatGPT to train itself on scraped content?

Re: Big Tech's underground race to buy AI training data

#82
post #75
post #70

Earlier quoted context omitted.

Google surfaces data — or it used to — LLMs and AI companies actively exploit it with zero benefit given to creators or users of the platforms they're now cannibalizing.

the irony. im surprised how businesses built on selling google search results is allowed to exist. i guess for the same reason google scraping the internet and building a product on top of it is allowed. then it only makes sense scraped AI training data is also going to be tolerated because you would need to reproduce a large language model like ChatGPT using your copyrighted content can produce a similar derivative…

The whole notion that AI can replace search is nonsense. It yields no benefit to the creators of the results it scrapes and the models hallucinate. It's worse for users and it's worse for everyone producing anything of note online.

Re: Big Tech's underground race to buy AI training data

#83
post #73
post #71

Earlier quoted context omitted.

If the predictions that traditional search engines will be displaced by LLM engines turn out to be correct then there will have to be a reckoning about copyright. It's already difficult enough to make money by writing online, but if most content gets consumed second-hand through an LLM then it will become basically impossible. How are journalists supposed to eat if NewsGPT just scoops up their work and starts regurgi…

How are you supposed to trust "journalism" from a text generator that hallucinates? The information ecosystem is bad enough without running it through a text blender that's already hitting compute, power and data limits.

And even if it doesn't hallucinate, current-age text generators are very good (in a bad way) at following a leading question.

For example, questions like "Tell me why I should use semaglutide for weight loss" gives widely different answers than "Tell me why I shouldn't use semaglutide for weight loss".

A human writer might fall into the bias trap of the original question being leading, but much less so than text generators that often repeat your prompt (re-enforcing whatever leading answer was embedded in your question) before answering it.

Re: Big Tech's underground race to buy AI training data

#84
post #81
post #77

Earlier quoted context omitted.

The same benefit doesn’t exist for ChatGPT as Google because Google means people click on your site and you get ad revenue. Google even facilitates this in both directions with search ads and as an ad service you can get paid from for hosting ads. The ROM site DMCA thing was always BS lmao it’s completely legal for you to dump your own carts and use them in emulators but that freedom doesn’t extend to having a copy o…

so you think scraping copyrighted content to sell ads is okay and downloading copyrighted games for free is also okay then why is it not okay for ChatGPT to train itself on scraped content?

It's not scraping, it's indexing and linking out to creators. LLMs are helping themselves to everything with no regard for content creators. They should be subject to copyright claims — I don't care if it destroys their business, they should've considered that at the outset. They didn't then and they don't care to now, they're simply greedy and looking to build something that benefits themselves and their investors with no regard for anyone they step on to do so.

Re: Big Tech's underground race to buy AI training data

#85
post #59
post #56

All the more reason for comprehensive privacy/data protection legislation and a refusal to provide data to these companies wherever possible.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

The main counterargument is that you have read 1000s of documents to train your brain which produces unique documents with no credit to the original copyright holders.

GenAI is just doing the same thing on a larger scale.

Re: Big Tech's underground race to buy AI training data

#86

Earlier quoted context omitted.

We can be forgiven for not having foreseen how social media would be used against us, connecting the world sounded like a cool idea on the surface. But having gone through that there's no excuse to be naive and simplistic about AI.

What harm has been inflicted upon you directly as a result?

That's like asking: "What harm has been inflicted upon you, individually and directly by some parts-per-million of a cancer-promoting chemical dumped in the town's water supply over the last twenty years?"

Re: Big Tech's underground race to buy AI training data

#87
post #60

I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.

This already happens. I have seen recruiters trying to get domain experts in various fields to write articles for AI training.

Re: Big Tech's underground race to buy AI training data

#88
post #81
post #77

Earlier quoted context omitted.

The same benefit doesn’t exist for ChatGPT as Google because Google means people click on your site and you get ad revenue. Google even facilitates this in both directions with search ads and as an ad service you can get paid from for hosting ads. The ROM site DMCA thing was always BS lmao it’s completely legal for you to dump your own carts and use them in emulators but that freedom doesn’t extend to having a copy o…

so you think scraping copyrighted content to sell ads is okay and downloading copyrighted games for free is also okay then why is it not okay for ChatGPT to train itself on scraped content?

The first part is fine because the search engine blurb isn’t a replacement for the thing itself. And I disagree with what ROM sites claim, you can’t just dump ROMs online and claim it’s not copyright infringement

Re: Big Tech's underground race to buy AI training data

#89
post #71
post #59

Earlier quoted context omitted.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

If the predictions that traditional search engines will be displaced by LLM engines turn out to be correct then there will have to be a reckoning about copyright. It's already difficult enough to make money by writing online, but if most content gets consumed second-hand through an LLM then it will become basically impossible. How are journalists supposed to eat if NewsGPT just scoops up their work and starts regurgi…

> How are journalists supposed to eat if NewsGPT just scoops up their work and starts regurgitating it seconds after publishing?

NewsGPT won't just regurgitate the work of journalists. First it'll consider the paid "partners" of NewsGPT to make sure to downplay anything that might hurt them, then it'll do the same for their advertisers while inserting some ads in the text, then they'll give the article tweaks according to NewsGPT's own ideology and then finally spit out something very different at their users. Maybe they can argue that NewsGPT is too transformative to count as copyright infringement.

Re: Big Tech's underground race to buy AI training data

#90

Who could have guessed giving away all of our data to corporations wholly focused on profit would be a bad thing?

If the end result is ai chat agents that anyone in the world can access for free, that seems like an absolutely wonderful thing

Are you an artist selling their own work?
Post reply on HN