All the more reason for comprehensive privacy/data protection legislation and a refusal to provide data to these companies wherever possible.
The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people
Big Tech's underground race to buy AI training data
71–80 of 152 posts
Re: Big Tech's underground race to buy AI training data
#72>in talks with multiple tech companies to license Photobucket's 13 billion photos and videos >Photobucket declined to identify its prospective buyers, citing commercial confidentiality. >tech companies are also quietly paying for content locked behind paywalls and login screens, giving rise to a hidden trade in everything from chat logs to long forgotten personal photos from faded social media apps In this market, et…
A lot of people just were not paying attention to the game being played, and so now they're getting played themselves.
Re: Big Tech's underground race to buy AI training data
#73Earlier quoted context omitted.
The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people
If the predictions that traditional search engines will be displaced by LLM engines turn out to be correct then there will have to be a reckoning about copyright. It's already difficult enough to make money by writing online, but if most content gets consumed second-hand through an LLM then it will become basically impossible. How are journalists supposed to eat if NewsGPT just scoops up their work and starts regurgi…
Re: Big Tech's underground race to buy AI training data
#74I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.
Re: Big Tech's underground race to buy AI training data
#75Earlier quoted context omitted.
isnt that what Google did ? they scraped the internet but the public/econ advisors felt the benefits outweighed copyright violations, they were just "indexers", they weren't scraping "news" they were indexing it lol same thing with emulators and roms. somebody dumped the cartridges (copyrighted software) into ROM files to be played on emulators (copyrighted bios) but they were "archiving" and if you owned the origina…
Google surfaces data — or it used to — LLMs and AI companies actively exploit it with zero benefit given to creators or users of the platforms they're now cannibalizing.
then it only makes sense scraped AI training data is also going to be tolerated because you would need to reproduce a large language model like ChatGPT using your copyrighted content can produce a similar derivative of your copyrighted content by doing forensic analysis.
its such an uphill battle for copyright holders. They need to replicate: copyrighted input ---> LM similar to ChatGPT4 ---> copyrighted output
So far its not looking good for OpenAI because its possible to generate copyrighted output (type spiderman in czech) so all that remains is demonstrating the middle layer (training it on LM similar to ChatGPT4) but that is unrealistically expensive.
I have theory that all this money spent on large models is to make it impossible for discovery (as it would require access to $100 billion GPUs)
Re: Big Tech's underground race to buy AI training data
#76I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.
Re: Big Tech's underground race to buy AI training data
#77Earlier quoted context omitted.
The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people
isnt that what Google did ? they scraped the internet but the public/econ advisors felt the benefits outweighed copyright violations, they were just "indexers", they weren't scraping "news" they were indexing it lol same thing with emulators and roms. somebody dumped the cartridges (copyrighted software) into ROM files to be played on emulators (copyrighted bios) but they were "archiving" and if you owned the origina…
Re: Big Tech's underground race to buy AI training data
#78Who could have guessed giving away all of our data to corporations wholly focused on profit would be a bad thing?
If the end result is ai chat agents that anyone in the world can access for free, that seems like an absolutely wonderful thing
Re: Big Tech's underground race to buy AI training data
#79Earlier quoted context omitted.
We can be forgiven for not having foreseen how social media would be used against us, connecting the world sounded like a cool idea on the surface. But having gone through that there's no excuse to be naive and simplistic about AI.
What harm has been inflicted upon you directly as a result?
Your health insurance company can buy up records from a data broker that show you've been spending 6% more time at fast food restaurants compared to last year and they can use that to raise your rates, but they'll never tell you that was the reason, you'll just have a higher bill than before
An employer can pass you over for a job you've applied to because you wrote something on a social media site 12 years ago that offended their political ideology, but you'll never know that was why, you'll just never get a call back.
If you get arrested, you could be denied bail because some AI decided you were a flight risk or more likely to reoffend if released, but no one will be able to tell you what made the AI decide that and you may not even be told an AI was used to make that choice.
As our lives become increasingly interconnected and recorded and analyzed it becomes extremely difficult for you to be aware of how or why your data is impacting your life, but it would be a huge mistake to assume that it isn't.
Re: Big Tech's underground race to buy AI training data
#80Earlier quoted context omitted.
The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people
If the predictions that traditional search engines will be displaced by LLM engines turn out to be correct then there will have to be a reckoning about copyright. It's already difficult enough to make money by writing online, but if most content gets consumed second-hand through an LLM then it will become basically impossible. How are journalists supposed to eat if NewsGPT just scoops up their work and starts regurgi…