Live data from Hacker News

Big Tech's underground race to buy AI training data

reuters.com

101–110 of 152 posts

Re: Big Tech's underground race to buy AI training data

#101
post #60

I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.

I have to imagine the valuable training data is domain specific stuff like sales call recordings for specific industries and technical materials about specific topics owned by companies. Surely there is enough public or copyright free general purpose material.

[deleted]

Re: Big Tech's underground race to buy AI training data

#102

This market is troubling. But I have a different question: What does the long game look like for raw training data? How will AIs maintain the quality of their diet? To compare, web search started — in the early days of Google — as a huge win because so much valuable information that was scattered around became findable. But over time it has become whac-a-mole with spam and AI copypasta, and now it's a struggle to kee…

Just like how ads have integrated into everything, trying to get us to click away from the happy path, AI will be in everything, trying to get us to do things that it is not yet good at so that it can learn from us. Which would be fine if the newfound efficiencies were properly democratized.

Yep. All these tech giants are taking the labour that people provided to the world in good faith in what I've seen described as a gift economy, and trying to lock it up. On the internet for the longest time people were providing their knowledge and fruits of labour for free, anticipating reciprocity (which on average they got). They stopped when reciprocity stopped. Platforms would monetize their efforts, control the distribution and often remove the reference to the creator.

These AI systems are being build on top of all the collective effort and resulting knowledge of the entire humanity. We can pretend they are just another private enterprise or we can acknowledge that they are something more than that.

And it's not just the productivity we could achieve with democratizing these systems. There's another danger. When big companies buy up all this intellectual property, what better choice would they have than to lock it up? At least until recently you could argue that IP rights owners were as entities incentivized to proliferate this knowledge, now the opposite is happening.

Re: Big Tech's underground race to buy AI training data

#103
post #2

>Rates vary by buyer and content type, but Braga said companies are generally willing to pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films. The market rate for text is $0.001 per word, she added. This is high enough that there should be a market to compensate the end users who created these

If you give something away when it's worthless, don't come back for more when it's discovered to be worth more.

Users of these sites have had license agreements and privacy policies for a long time, and freely gave away their content just because free web hosting was worth it. Why would they be entitled to anything more now that this content have found new value?

Re: Big Tech's underground race to buy AI training data

#104
post #93

Earlier quoted context omitted.

Regurgitation seconds after is what already happens with the AP though. There are some real journalists that will sadly be pushed further out of the fold, and presumably many human but fake journalists that have been coasting for years on such regurgitation. I’m not so optimistic about the ai future, and believe payment or at least credit really needs to get figured out for generative stuff. But real content producer…

> Regurgitation seconds after is what already happens with the AP though. The AP makes about half a billion a year from other outlets paying them for permission to regurgitate their content. That's not the same as the AI lobby saying they should be allowed to scrape apnews.com and publish articles derived from the content they get from there, for free, and without attribution.

I do see your point, and yeah theft is theft and theft is bad. But what I’m getting at is more the POV for media consumers.

If it’s regurgitated / unoriginal anyway then I don’t think most people care much whether it’s summarized/subjected to extra spin and fluff by a person or by a machine.

Re: Big Tech's underground race to buy AI training data

#105

Earlier quoted context omitted.

If the end result is ai chat agents that anyone in the world can access for free, that seems like an absolutely wonderful thing

Are you an artist selling their own work?

Code. Every day.

When you solve something tricky, you just basically released that as open source and trained ChatGPT 2030 how to do it without you.

Re: Big Tech's underground race to buy AI training data

#106
post #84
post #81

Earlier quoted context omitted.

so you think scraping copyrighted content to sell ads is okay and downloading copyrighted games for free is also okay then why is it not okay for ChatGPT to train itself on scraped content?

It's not scraping, it's indexing and linking out to creators. LLMs are helping themselves to everything with no regard for content creators. They should be subject to copyright claims — I don't care if it destroys their business, they should've considered that at the outset. They didn't then and they don't care to now, they're simply greedy and looking to build something that benefits themselves and their investors w…

but how can you prove that your picture of a cat was used in LLM?

if you owned a franchise called "Chicken Brothers" with a the logo of two chickens standing side by side with arms crossed proudly then do you have claim over all derivatives including the spanish name generated by LLM?

i just dont think its straight forward, the main complaint should be payout for license used during training but its tough to prove unless someone at OpenAI dumps the AWS cloudwatch logs

Re: Big Tech's underground race to buy AI training data

#107
post #82
post #75

Earlier quoted context omitted.

the irony. im surprised how businesses built on selling google search results is allowed to exist. i guess for the same reason google scraping the internet and building a product on top of it is allowed. then it only makes sense scraped AI training data is also going to be tolerated because you would need to reproduce a large language model like ChatGPT using your copyrighted content can produce a similar derivative…

The whole notion that AI can replace search is nonsense. It yields no benefit to the creators of the results it scrapes and the models hallucinate. It's worse for users and it's worse for everyone producing anything of note online.

but many chatgpt users are not using Google as much instead relying on LLMs + RAG

ChatGPT is the new search engine and provides far more value to the end user than Google.

The issue seems to be people want a payout from OpenAI...but its non-profit

Re: Big Tech's underground race to buy AI training data

#108
post #17

Earlier quoted context omitted.

If the end result is ai chat agents that anyone in the world can access for free, that seems like an absolutely wonderful thing

If the companies making those agents are paying top dollar for training data then the product isn't going to be truly free, at best it will be "free" with caveats. Do you want to use an AI agent which is fine-tuned according to the wishes of the top bidding advertisers? Because that's probably the first thing they'll try to make "free" chatbots actually turn a profit. The future is having your own personal AI assista…

As Yuval Harari suggests the AI economy will move away from money and man power. What will be important are control over resources and their distribution. These big companies won't care about a number in some database. They won't care about selling stuff to you, maybe in the midterm but not in the long run.

Re: Big Tech's underground race to buy AI training data

#109
post #65

Earlier quoted context omitted.

I love this fact! I would have never realized it

These are probably pretty arbitrarily priced so someone thought they'd be cute & everyone else picked up this rate.

I don't think market prices usually work this way. Am I missing a reason that this is an exception?

Re: Big Tech's underground race to buy AI training data

#110
post #60

I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.

This won't be necessary in future AIs. As AIs will start aligning tokens from all the rich modalities of audio, video, 3D with text so that they can express complex ideas, they will bootstrap in proper language generation.

I don't think college essays, etc would contain anything novel. Future techniques could smoothly interpolate better creating ever-anew wordmud.

Post reply on HN