I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.
I have to imagine the valuable training data is domain specific stuff like sales call recordings for specific industries and technical materials about specific topics owned by companies. Surely there is enough public or copyright free general purpose material.
Big Tech's underground race to buy AI training data
101–110 of 152 posts
Re: Big Tech's underground race to buy AI training data
#102This market is troubling. But I have a different question: What does the long game look like for raw training data? How will AIs maintain the quality of their diet? To compare, web search started — in the early days of Google — as a huge win because so much valuable information that was scattered around became findable. But over time it has become whac-a-mole with spam and AI copypasta, and now it's a struggle to kee…
Just like how ads have integrated into everything, trying to get us to click away from the happy path, AI will be in everything, trying to get us to do things that it is not yet good at so that it can learn from us. Which would be fine if the newfound efficiencies were properly democratized.
These AI systems are being build on top of all the collective effort and resulting knowledge of the entire humanity. We can pretend they are just another private enterprise or we can acknowledge that they are something more than that.
And it's not just the productivity we could achieve with democratizing these systems. There's another danger. When big companies buy up all this intellectual property, what better choice would they have than to lock it up? At least until recently you could argue that IP rights owners were as entities incentivized to proliferate this knowledge, now the opposite is happening.
Re: Big Tech's underground race to buy AI training data
#103>Rates vary by buyer and content type, but Braga said companies are generally willing to pay $1 to $2 per image, $2 to $4 per short-form video and $100 to $300 per hour of longer films. The market rate for text is $0.001 per word, she added. This is high enough that there should be a market to compensate the end users who created these
Users of these sites have had license agreements and privacy policies for a long time, and freely gave away their content just because free web hosting was worth it. Why would they be entitled to anything more now that this content have found new value?
Re: Big Tech's underground race to buy AI training data
#104Earlier quoted context omitted.
Regurgitation seconds after is what already happens with the AP though. There are some real journalists that will sadly be pushed further out of the fold, and presumably many human but fake journalists that have been coasting for years on such regurgitation. I’m not so optimistic about the ai future, and believe payment or at least credit really needs to get figured out for generative stuff. But real content producer…
> Regurgitation seconds after is what already happens with the AP though. The AP makes about half a billion a year from other outlets paying them for permission to regurgitate their content. That's not the same as the AI lobby saying they should be allowed to scrape apnews.com and publish articles derived from the content they get from there, for free, and without attribution.
If it’s regurgitated / unoriginal anyway then I don’t think most people care much whether it’s summarized/subjected to extra spin and fluff by a person or by a machine.
Re: Big Tech's underground race to buy AI training data
#105Earlier quoted context omitted.
If the end result is ai chat agents that anyone in the world can access for free, that seems like an absolutely wonderful thing
Are you an artist selling their own work?
When you solve something tricky, you just basically released that as open source and trained ChatGPT 2030 how to do it without you.
Re: Big Tech's underground race to buy AI training data
#106Earlier quoted context omitted.
so you think scraping copyrighted content to sell ads is okay and downloading copyrighted games for free is also okay then why is it not okay for ChatGPT to train itself on scraped content?
It's not scraping, it's indexing and linking out to creators. LLMs are helping themselves to everything with no regard for content creators. They should be subject to copyright claims — I don't care if it destroys their business, they should've considered that at the outset. They didn't then and they don't care to now, they're simply greedy and looking to build something that benefits themselves and their investors w…
if you owned a franchise called "Chicken Brothers" with a the logo of two chickens standing side by side with arms crossed proudly then do you have claim over all derivatives including the spanish name generated by LLM?
i just dont think its straight forward, the main complaint should be payout for license used during training but its tough to prove unless someone at OpenAI dumps the AWS cloudwatch logs
Re: Big Tech's underground race to buy AI training data
#107Earlier quoted context omitted.
the irony. im surprised how businesses built on selling google search results is allowed to exist. i guess for the same reason google scraping the internet and building a product on top of it is allowed. then it only makes sense scraped AI training data is also going to be tolerated because you would need to reproduce a large language model like ChatGPT using your copyrighted content can produce a similar derivative…
The whole notion that AI can replace search is nonsense. It yields no benefit to the creators of the results it scrapes and the models hallucinate. It's worse for users and it's worse for everyone producing anything of note online.
ChatGPT is the new search engine and provides far more value to the end user than Google.
The issue seems to be people want a payout from OpenAI...but its non-profit
Re: Big Tech's underground race to buy AI training data
#108Earlier quoted context omitted.
If the end result is ai chat agents that anyone in the world can access for free, that seems like an absolutely wonderful thing
If the companies making those agents are paying top dollar for training data then the product isn't going to be truly free, at best it will be "free" with caveats. Do you want to use an AI agent which is fine-tuned according to the wishes of the top bidding advertisers? Because that's probably the first thing they'll try to make "free" chatbots actually turn a profit. The future is having your own personal AI assista…
Re: Big Tech's underground race to buy AI training data
#109Earlier quoted context omitted.
I love this fact! I would have never realized it
These are probably pretty arbitrarily priced so someone thought they'd be cute & everyone else picked up this rate.
Re: Big Tech's underground race to buy AI training data
#110I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.
I don't think college essays, etc would contain anything novel. Future techniques could smoothly interpolate better creating ever-anew wordmud.