Live data from Hacker News

Big Tech's underground race to buy AI training data

reuters.com

61–70 of 152 posts

Re: Big Tech's underground race to buy AI training data

#61

Who could have guessed giving away all of our data to corporations wholly focused on profit would be a bad thing?

If the end result is ai chat agents that anyone in the world can access for free, that seems like an absolutely wonderful thing

> for free

Even if a future service doesn't have an obvious charge or subscription, just because you don't recognize how you're being exploited doesn't mean it's truly "free."

There's a reason advertising exists as an industry at all, let alone a global trillion-dollar one. Today's "free" is actually paid for by exploiting user attention and attempting to hack your brain--sometimes in ways that are culturally accepted due to long tradition of use, sometimes in new disturbing ones.

Re: Big Tech's underground race to buy AI training data

#62

Earlier quoted context omitted.

We can be forgiven for not having foreseen how social media would be used against us, connecting the world sounded like a cool idea on the surface. But having gone through that there's no excuse to be naive and simplistic about AI.

What harm has been inflicted upon you directly as a result?

I imagine that probably /can/ name direct harms but why does someone need to have been directly harmed for them to be concerned about it/think it's bad? I've never been murdered but that's hardly reason for me to be okay with murder.

Re: Big Tech's underground race to buy AI training data

#63
post #5

They talk about voice samples, but they don’t mention prices for them Would it be attractive for a company like Twilio or Aircall to offer free phone calls and sell anonymized recordings?

Funnily, this is how Google improved their voice recognition. Remember a decade or so ago, you could call a 1-800 number and look up phone numbers using your voice? It was backed by Google and once Google was done collecting the data, they shut it down.

Google411

Re: Big Tech's underground race to buy AI training data

#64
post #60

I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.

Most companies are hiring for the role of AI Tutor. Some of that is definitely happening.

Re: Big Tech's underground race to buy AI training data

#65
post #2

>Rates vary by buyer and content type, but Braga said companies are generally willing to pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films. The market rate for text is $0.001 per word, she added. This is high enough that there should be a market to compensate the end users who created these

> The market rate for text is $0.001 per word, she added. I'm astonished that a picture turns out to be worth a thousand words.

I love this fact! I would have never realized it

Re: Big Tech's underground race to buy AI training data

#66
post #59
post #56

All the more reason for comprehensive privacy/data protection legislation and a refusal to provide data to these companies wherever possible.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

isnt that what Google did ? they scraped the internet but the public/econ advisors felt the benefits outweighed copyright violations, they were just "indexers", they weren't scraping "news" they were indexing it lol

same thing with emulators and roms. somebody dumped the cartridges (copyrighted software) into ROM files to be played on emulators (copyrighted bios) but they were "archiving" and if you owned the original copy you could download them. I still vividly remember seeing on warez website disclaimer: "DMCA SAFE HARBOUR NOTICE: YOU MUST OWN THE ORIGINAL GAME OTHERWISE ITS ILLEGAL BUT YES, YOU CAN DOWNLOAD EVERY SINGLE GAME MADE ON THAT CONSOLE FOR FREE"

I feel like the same outcome will be for LLMs trained on copyrighted material. It will be "training". The net benefit is too great than fretting over "training"

tldr: "indexing" ---> "archiving" ---> "training"

Re: Big Tech's underground race to buy AI training data

#67
post #24

Earlier quoted context omitted.

No, that's gross violation of privacy; no such thing as anonymized recordings.

It would be a violation of privacy if people weren’t aware/hadn’t consented But if it was part of the terms of the new free service, and all the parties involved got a reminder message on the call… you might still not like it, but it doesn’t seem like it would be a violation of privacy

A fantastic example of why the inherent "lack" of one party in an economic exchange is a necessary component of modern capitalism.

The only people who would be willing to use such a service are people who have likely already been systematically disenfranchised by our global economic system. Poor people.

Privacy should not be incentivized and treated as a luxury. Especially when the end result of all this training data is models which further discriminate against vulnerable third-parties and automate maximum value extraction from the average user via unprecedented amounts of emotional manipulation afforded to us by the development of user-facing generative AI. Whether through highly-targeted, ad-hoc advertisements, or discriminative insurance policies.

Re: Big Tech's underground race to buy AI training data

#68
post #65

Earlier quoted context omitted.

> The market rate for text is $0.001 per word, she added. I'm astonished that a picture turns out to be worth a thousand words.

I love this fact! I would have never realized it

These are probably pretty arbitrarily priced so someone thought they'd be cute & everyone else picked up this rate.

Re: Big Tech's underground race to buy AI training data

#69
post #59
post #56

All the more reason for comprehensive privacy/data protection legislation and a refusal to provide data to these companies wherever possible.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

Frankly, it is and should be treated as such. The fact that they're dodging questions about their data sources is a red flag and a pretty clear indication that they know they're in the wrong and are fighting to become established enough to be in a position to, at best, ask for forgiveness after the fact.

Re: Big Tech's underground race to buy AI training data

#70
post #66
post #59

Earlier quoted context omitted.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

isnt that what Google did ? they scraped the internet but the public/econ advisors felt the benefits outweighed copyright violations, they were just "indexers", they weren't scraping "news" they were indexing it lol same thing with emulators and roms. somebody dumped the cartridges (copyrighted software) into ROM files to be played on emulators (copyrighted bios) but they were "archiving" and if you owned the origina…

Google surfaces data — or it used to — LLMs and AI companies actively exploit it with zero benefit given to creators or users of the platforms they're now cannibalizing.
Post reply on HN