Live data from Hacker News

Big Tech's underground race to buy AI training data

reuters.com

71–80 of 152 posts

Re: Big Tech's underground race to buy AI training data

#71
post #59
post #56

All the more reason for comprehensive privacy/data protection legislation and a refusal to provide data to these companies wherever possible.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

If the predictions that traditional search engines will be displaced by LLM engines turn out to be correct then there will have to be a reckoning about copyright. It's already difficult enough to make money by writing online, but if most content gets consumed second-hand through an LLM then it will become basically impossible. How are journalists supposed to eat if NewsGPT just scoops up their work and starts regurgitating it seconds after publishing?

Re: Big Tech's underground race to buy AI training data

#72

>in talks with multiple tech companies to license Photobucket's 13 billion photos and videos >Photobucket declined to identify its prospective buyers, citing commercial confidentiality. >tech companies are also quietly paying for content locked behind paywalls and login screens, giving rise to a hidden trade in everything from chat logs to long forgotten personal photos from faded social media apps In this market, et…

I vividly remember consenting to all variety of terms agreements as a 13-year old on the web in 2007. I also remember explicitly licensing all of my output as CC and embracing copyleft. It's never been a secret that even captchas contribute to the improvement of models designed to ultimately sell ads to eyeballs.

A lot of people just were not paying attention to the game being played, and so now they're getting played themselves.

Re: Big Tech's underground race to buy AI training data

#73
post #71
post #59

Earlier quoted context omitted.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

If the predictions that traditional search engines will be displaced by LLM engines turn out to be correct then there will have to be a reckoning about copyright. It's already difficult enough to make money by writing online, but if most content gets consumed second-hand through an LLM then it will become basically impossible. How are journalists supposed to eat if NewsGPT just scoops up their work and starts regurgi…

How are you supposed to trust "journalism" from a text generator that hallucinates? The information ecosystem is bad enough without running it through a text blender that's already hitting compute, power and data limits.

Re: Big Tech's underground race to buy AI training data

#74
post #60

I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.

I have to imagine the valuable training data is domain specific stuff like sales call recordings for specific industries and technical materials about specific topics owned by companies. Surely there is enough public or copyright free general purpose material.

Re: Big Tech's underground race to buy AI training data

#75
post #70
post #66

Earlier quoted context omitted.

isnt that what Google did ? they scraped the internet but the public/econ advisors felt the benefits outweighed copyright violations, they were just "indexers", they weren't scraping "news" they were indexing it lol same thing with emulators and roms. somebody dumped the cartridges (copyrighted software) into ROM files to be played on emulators (copyrighted bios) but they were "archiving" and if you owned the origina…

Google surfaces data — or it used to — LLMs and AI companies actively exploit it with zero benefit given to creators or users of the platforms they're now cannibalizing.

the irony. im surprised how businesses built on selling google search results is allowed to exist. i guess for the same reason google scraping the internet and building a product on top of it is allowed.

then it only makes sense scraped AI training data is also going to be tolerated because you would need to reproduce a large language model like ChatGPT using your copyrighted content can produce a similar derivative of your copyrighted content by doing forensic analysis.

its such an uphill battle for copyright holders. They need to replicate: copyrighted input ---> LM similar to ChatGPT4 ---> copyrighted output

So far its not looking good for OpenAI because its possible to generate copyrighted output (type spiderman in czech) so all that remains is demonstrating the middle layer (training it on LM similar to ChatGPT4) but that is unrealistically expensive.

I have theory that all this money spent on large models is to make it impossible for discovery (as it would require access to $100 billion GPUs)

Re: Big Tech's underground race to buy AI training data

#76
post #60

I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.

They're more interested in eliminating jobs than creating them.

Re: Big Tech's underground race to buy AI training data

#77
post #66
post #59

Earlier quoted context omitted.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

isnt that what Google did ? they scraped the internet but the public/econ advisors felt the benefits outweighed copyright violations, they were just "indexers", they weren't scraping "news" they were indexing it lol same thing with emulators and roms. somebody dumped the cartridges (copyrighted software) into ROM files to be played on emulators (copyrighted bios) but they were "archiving" and if you owned the origina…

The same benefit doesn’t exist for ChatGPT as Google because Google means people click on your site and you get ad revenue. Google even facilitates this in both directions with search ads and as an ad service you can get paid from for hosting ads. The ROM site DMCA thing was always BS lmao it’s completely legal for you to dump your own carts and use them in emulators but that freedom doesn’t extend to having a copy of someone else’s game cart. That’s just an intentional misunderstanding of the DMCA in a futile attempt to not get banned

Re: Big Tech's underground race to buy AI training data

#78

Who could have guessed giving away all of our data to corporations wholly focused on profit would be a bad thing?

If the end result is ai chat agents that anyone in the world can access for free, that seems like an absolutely wonderful thing

Perplexity is already looking to include ads. Free is a hook to get users invested before they trap them and extract all the value for themselves.

Re: Big Tech's underground race to buy AI training data

#79

Earlier quoted context omitted.

We can be forgiven for not having foreseen how social media would be used against us, connecting the world sounded like a cool idea on the surface. But having gone through that there's no excuse to be naive and simplistic about AI.

What harm has been inflicted upon you directly as a result?

That's the beauty of the system we have. Your data can be routinely used against you and you'll never be told about it.

Your health insurance company can buy up records from a data broker that show you've been spending 6% more time at fast food restaurants compared to last year and they can use that to raise your rates, but they'll never tell you that was the reason, you'll just have a higher bill than before

An employer can pass you over for a job you've applied to because you wrote something on a social media site 12 years ago that offended their political ideology, but you'll never know that was why, you'll just never get a call back.

If you get arrested, you could be denied bail because some AI decided you were a flight risk or more likely to reoffend if released, but no one will be able to tell you what made the AI decide that and you may not even be told an AI was used to make that choice.

As our lives become increasingly interconnected and recorded and analyzed it becomes extremely difficult for you to be aware of how or why your data is impacting your life, but it would be a huge mistake to assume that it isn't.

Re: Big Tech's underground race to buy AI training data

#80
post #71
post #59

Earlier quoted context omitted.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

If the predictions that traditional search engines will be displaced by LLM engines turn out to be correct then there will have to be a reckoning about copyright. It's already difficult enough to make money by writing online, but if most content gets consumed second-hand through an LLM then it will become basically impossible. How are journalists supposed to eat if NewsGPT just scoops up their work and starts regurgi…

Regurgitation seconds after is what already happens with the AP though. There are some real journalists that will sadly be pushed further out of the fold, and presumably many human but fake journalists that have been coasting for years on such regurgitation. I’m not so optimistic about the ai future, and believe payment or at least credit really needs to get figured out for generative stuff. But real content producers should direct some of the irritation at their editors, colleagues, and industry or else it’s all rather dishonest isn’t it?
Post reply on HN