Live data from Hacker News

Big Tech's underground race to buy AI training data

reuters.com

41–50 of 152 posts

Re: Big Tech's underground race to buy AI training data

#42
post #2

>Rates vary by buyer and content type, but Braga said companies are generally willing to pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films. The market rate for text is $0.001 per word, she added. This is high enough that there should be a market to compensate the end users who created these

Are counterfeit words then AI generated? Just like money you need a very good “press” and hard to detect..

Re: Big Tech's underground race to buy AI training data

#45
post #2

>Rates vary by buyer and content type, but Braga said companies are generally willing to pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films. The market rate for text is $0.001 per word, she added. This is high enough that there should be a market to compensate the end users who created these

Are certain types of textual content more valuable than others? For instance, conversations vs long form content vs short form (ie tweets)

Re: Big Tech's underground race to buy AI training data

#46

Google having so many private photos in Google Photos must be a goldmine for them.

> Google having so many private photos in Google Photos must be a goldmine for them. While true, it's META who has won that arm's race long ago in my view; hell, they just disclosed that they have private access to DMs to Netflixh [0] in a lawsuit. If you don;t think they are training their own models on this data over all their platforms you have to be a complete idiot o: Facebook, Instagram, Whatsapp. That is a muc…

Whatsapp chats are encrypted, how can they be used to train the models? Also what kind of training can be done on Instagram data, is there anything of value there?

Re: Big Tech's underground race to buy AI training data

#47
post #6

This will be a fun reminiscence once we find out how humans are able to learn with just a tiny fraction of that data volume.

Tiny fraction... if you ignore the learning data processed by a billion years of evolution.

It’s a good question what portion of our DNA contributes to the information processing and knowledge in our brain.

However, the first complex nervous systems came about in the Cambrian explosion, only about half a billion years ago. And we also don’t train LLMs by random mutation and selection, it’s a much more teleological process.

But to extend the analogy, we should be able to train a model continuously, and not have to start training from scratch for each new model. Although, maybe, that would require random mutations, and thus much more time?

Re: Big Tech's underground race to buy AI training data

#49
post #35

I assume some of the more shady/no-name dashcam units with Wifi capability are uploading their video and internal microphone recordings. Distributed surveillance: The Panopitcar

I've wondered about crowdsourcing that. Sousveillance. Don't think enough people would be interested, though.

Re: Big Tech's underground race to buy AI training data

#50

Google having so many private photos in Google Photos must be a goldmine for them.

> Google having so many private photos in Google Photos must be a goldmine for them. While true, it's META who has won that arm's race long ago in my view; hell, they just disclosed that they have private access to DMs to Netflixh [0] in a lawsuit. If you don;t think they are training their own models on this data over all their platforms you have to be a complete idiot o: Facebook, Instagram, Whatsapp. That is a muc…

> Google is limited to mainly Android users

https://www.appmysite.com/blog/android-vs-ios-mobile-operati...

Random link. Can't vouch for it. But US and RoW have quite different patterns.

Post reply on HN