Live data from Hacker News

Big Tech's underground race to buy AI training data

reuters.com

91–100 of 152 posts

Re: Big Tech's underground race to buy AI training data

#91
post #85
post #59

Earlier quoted context omitted.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

The main counterargument is that you have read 1000s of documents to train your brain which produces unique documents with no credit to the original copyright holders. GenAI is just doing the same thing on a larger scale.

Well, if the concept is "we should legally treat the systems just like very big humans", then the next step is to arrest and confine all the leaders of the companies involved on charges of slavery and child exploitation.

The distinction does matter in copyright too, since a transformative work needs some non-trivial amount of human input.

Re: Big Tech's underground race to buy AI training data

#92
post #53

>in talks with multiple tech companies to license Photobucket's 13 billion photos and videos >Photobucket declined to identify its prospective buyers, citing commercial confidentiality. >tech companies are also quietly paying for content locked behind paywalls and login screens, giving rise to a hidden trade in everything from chat logs to long forgotten personal photos from faded social media apps In this market, et…

It's immediately and self-evidently obvious that no end-user in 2007 consented to photos of their 2007 era teenage self being used to train an AI how to identify an emo kid Unfortunately, they did actually. It's more accurate to say that they were presented a EULA and Terms of Service that no reasonable teenager would have had any hope of understanding. But since they're over 13, they're held to the terms of those ag…

There's a funny thing where the legal/commercial definition of "consent" is essentially a subset of "non-consent", having extremely little overlap with "consent" in a meaningful way.

Re: Big Tech's underground race to buy AI training data

#93
post #71

Earlier quoted context omitted.

If the predictions that traditional search engines will be displaced by LLM engines turn out to be correct then there will have to be a reckoning about copyright. It's already difficult enough to make money by writing online, but if most content gets consumed second-hand through an LLM then it will become basically impossible. How are journalists supposed to eat if NewsGPT just scoops up their work and starts regurgi…

Regurgitation seconds after is what already happens with the AP though. There are some real journalists that will sadly be pushed further out of the fold, and presumably many human but fake journalists that have been coasting for years on such regurgitation. I’m not so optimistic about the ai future, and believe payment or at least credit really needs to get figured out for generative stuff. But real content producer…

> Regurgitation seconds after is what already happens with the AP though.

The AP makes about half a billion a year from other outlets paying them for permission to regurgitate their content. That's not the same as the AI lobby saying they should be allowed to scrape apnews.com and publish articles derived from the content they get from there, for free, and without attribution.

Re: Big Tech's underground race to buy AI training data

#94
post #71

Earlier quoted context omitted.

If the predictions that traditional search engines will be displaced by LLM engines turn out to be correct then there will have to be a reckoning about copyright. It's already difficult enough to make money by writing online, but if most content gets consumed second-hand through an LLM then it will become basically impossible. How are journalists supposed to eat if NewsGPT just scoops up their work and starts regurgi…

Regurgitation seconds after is what already happens with the AP though. There are some real journalists that will sadly be pushed further out of the fold, and presumably many human but fake journalists that have been coasting for years on such regurgitation. I’m not so optimistic about the ai future, and believe payment or at least credit really needs to get figured out for generative stuff. But real content producer…

We should be working to strengthen journalism as a practice. LLMs will do anything but.

Re: Big Tech's underground race to buy AI training data

#95
post #53

Earlier quoted context omitted.

It's immediately and self-evidently obvious that no end-user in 2007 consented to photos of their 2007 era teenage self being used to train an AI how to identify an emo kid Unfortunately, they did actually. It's more accurate to say that they were presented a EULA and Terms of Service that no reasonable teenager would have had any hope of understanding. But since they're over 13, they're held to the terms of those ag…

People have known for eons that companies were using their data. TikTok is well publicized to be the CCP. And yet millions of people would rather have entertainment. There are plenty (myself included) who abstain, but the reality is the vast majority of people, if presented with free and unlimited dopamine hits, will gladly give away their info.

> the reality is the vast majority of people, if presented with free and unlimited dopamine hits, will gladly give away their info

I counter that the reality is the vast majority of people do not meaningfully understand the exchange they are making. I'm not saying they're stupid or blaming them whatsoever; it's a similar phenomenon to playing the lottery. Our brains aren't equipped to understand such unintuitive phenomena.

Re: Big Tech's underground race to buy AI training data

#96

Earlier quoted context omitted.

> Google having so many private photos in Google Photos must be a goldmine for them. While true, it's META who has won that arm's race long ago in my view; hell, they just disclosed that they have private access to DMs to Netflixh [0] in a lawsuit. If you don;t think they are training their own models on this data over all their platforms you have to be a complete idiot o: Facebook, Instagram, Whatsapp. That is a muc…

Whatsapp chats are encrypted, how can they be used to train the models? Also what kind of training can be done on Instagram data, is there anything of value there?

> Also what kind of training can be done on Instagram data, is there anything of value there?

Billions of comments and private messages; billions of data points on user behavior and (more importantly) how they respond to manipulative UI/UX/content... Nothing useful there??

Re: Big Tech's underground race to buy AI training data

#97
post #35

I assume some of the more shady/no-name dashcam units with Wifi capability are uploading their video and internal microphone recordings. Distributed surveillance: The Panopitcar

Any modern car is likely to already be transmitting that data and more, such as your weight, metadata about your doctor visits, etc. Cars are a privacy nightmare.

Re: Big Tech's underground race to buy AI training data

#99
post #60

I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.

This already happens. I have seen recruiters trying to get domain experts in various fields to write articles for AI training.

LinkedIn built a whole platform inside their platform for doing exactly this. I think you get a badge or something on your profile claiming your an expert on something if you write a couple paragraphs on a topic using the provided prompt.

They're very clear its going into an AI generated article on the topic but you better believe that is also now core training data.

Re: Big Tech's underground race to buy AI training data

#100
post #2

>Rates vary by buyer and content type, but Braga said companies are generally willing to pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films. The market rate for text is $0.001 per word, she added. This is high enough that there should be a market to compensate the end users who created these

> The market rate for text is $0.001 per word, she added. I'm astonished that a picture turns out to be worth a thousand words.

On the other hand, per byte the word is more expensive
Post reply on HN