Live data from Hacker News

Big Tech's underground race to buy AI training data

reuters.com

141–150 of 152 posts

Re: Big Tech's underground race to buy AI training data

#142
post #85
post #59

Earlier quoted context omitted.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

The main counterargument is that you have read 1000s of documents to train your brain which produces unique documents with no credit to the original copyright holders. GenAI is just doing the same thing on a larger scale.

If I offered a paid service where you could pay me $20 a month and I would draw you copyrighted works that are in my internal neural network that would also be illegal

Re: Big Tech's underground race to buy AI training data

#143

Earlier quoted context omitted.

I vividly remember consenting to all variety of terms agreements as a 13-year old on the web in 2007. I also remember explicitly licensing all of my output as CC and embracing copyleft. It's never been a secret that even captchas contribute to the improvement of models designed to ultimately sell ads to eyeballs. A lot of people just were not paying attention to the game being played, and so now they're getting playe…

When a company hides their skeevy practices in 30-page social media term consent form, I blame them a lot more than the normal person with limited time or the literal child. I’d prefer such people didn’t “get played” by multinational corporations, even if they potentially could have prevented it.

I think it's important to understand that consent was indeed given, and most users likely understood that they did not own any non-copyrightable portion of their user-generated content.

Rather, the conversation should focus on how to improve parsing of ToS (I personally believe we should use symbolic labeling like we do with food), as well as regulation around what terms can change for content which was generated under the premise of an older ToS.

OP's statement, "It's immediately and self-evidently obvious that no end-user in 2007 consented to photos of their 2007 era teenage self being used to train an AI how to identify an emo kid," is simply false. Many, if not most users, understood that they gave permission for their UGC to be used to improve the services. This is what I am rebuking.

Re: Big Tech's underground race to buy AI training data

#144
post #71
post #59

Earlier quoted context omitted.

The fact that ChatGPT isn’t deemed copyright infringement is absurd. Like you can’t take the entire internet and use it to train your software and claim you’re not violating the copyright of thousands of people

If the predictions that traditional search engines will be displaced by LLM engines turn out to be correct then there will have to be a reckoning about copyright. It's already difficult enough to make money by writing online, but if most content gets consumed second-hand through an LLM then it will become basically impossible. How are journalists supposed to eat if NewsGPT just scoops up their work and starts regurgi…

Journalist are ultimately extremely overrated in 2024.

I go out of my way to not consume news outside what happens to cross my way because of financial markets.

What exactly do you think I am missing that is so important? Journalist by large produce complete nonsense in 2024. Journalist in 2024 are a massive net negative and would be much better served doing something productive, like selling apples on the street.

Re: Big Tech's underground race to buy AI training data

#145
post #100

Earlier quoted context omitted.

> The market rate for text is $0.001 per word, she added. I'm astonished that a picture turns out to be worth a thousand words.

On the other hand, per byte the word is more expensive

I’m an economist. This is an example of a volume discount. Prices often decrease per-unit when buying larger quantities. That happens whether it’s milk or the square-footage of an apartment. I’d expect larger files to be worth less per-byte. Photo and video files tend to be larger than text ones.

Re: Big Tech's underground race to buy AI training data

#146

Earlier quoted context omitted.

Photobucket is a morally bankrupt shell of its former self. They send constant emails with extremely urgent subject lines threatening to delete your photos unless you sign up for a $5/mo plan. They do this even if your account doesn't contain any photos .

> threatening to delete your photos unless you sign up for a $5/mo plan What's morally bankrupt about that? It costs money to host your photos and they're a business that can decide to charge their customers any rate they think the market will accept.

I have no photos on their service and they've been emailing me weekly since last year with URGENT and ACTION REQUIRED.

Re: Big Tech's underground race to buy AI training data

#147

Earlier quoted context omitted.

People have known for eons that companies were using their data. TikTok is well publicized to be the CCP. And yet millions of people would rather have entertainment. There are plenty (myself included) who abstain, but the reality is the vast majority of people, if presented with free and unlimited dopamine hits, will gladly give away their info.

Do you consider such people to blame? Under your own framing, they’re more or less being taken advantage of by a modern form of drug dealer.

Absolutely. Do you think drug addicts have no responsibility for their behavior? Are drug dealers viable without a strong customer base?

Re: Big Tech's underground race to buy AI training data

#148
post #95

Earlier quoted context omitted.

People have known for eons that companies were using their data. TikTok is well publicized to be the CCP. And yet millions of people would rather have entertainment. There are plenty (myself included) who abstain, but the reality is the vast majority of people, if presented with free and unlimited dopamine hits, will gladly give away their info.

> the reality is the vast majority of people, if presented with free and unlimited dopamine hits, will gladly give away their info I counter that the reality is the vast majority of people do not meaningfully understand the exchange they are making. I'm not saying they're stupid or blaming them whatsoever; it's a similar phenomenon to playing the lottery. Our brains aren't equipped to understand such unintuitive phen…

> I counter that the reality is the vast majority of people do not meaningfully understand the exchange they are making.

I’m not sure this is true in 2024 at all. The presumption today is you’re being tracked, and people simply don’t care.

But let’s presume this isn’t true. I think the response should be to expect more from society. Every additional bit of nanny state coddling reduces individual responsibility.

Re: Big Tech's underground race to buy AI training data

#149
post #60

I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.

People will just use ai to write those essays and emails!

Re: Big Tech's underground race to buy AI training data

#150
post #32

Earlier quoted context omitted.

The data volume is actually not that different once you account for all senses and how many years it takes for a human to become useful. The interesting thing would be how the human brain filters out the unimportant information as it develops.

That's a distinction without a difference. The majority of data is from a distribution that's already been sampled multiple times. E.g. how often does a baby go out and experience something novel? The majority of it's time is spent getting the same stimulus over and over again, as anyone listening to childrens television can attest. Humans learn in fundamentally different ways to our current systems and information p…

And what do you think epochs in machine learning are? Or why more modern training efforts (i.e. for LLMs) are focussing hard on deduplicating scraped data?
Post reply on HN