Live data from Hacker News

Big Tech's underground race to buy AI training data

reuters.com

131–140 of 152 posts

Re: Big Tech's underground race to buy AI training data

#131
post #96

Earlier quoted context omitted.

Whatsapp chats are encrypted, how can they be used to train the models? Also what kind of training can be done on Instagram data, is there anything of value there?

> Also what kind of training can be done on Instagram data, is there anything of value there? Billions of comments and private messages; billions of data points on user behavior and (more importantly) how they respond to manipulative UI/UX/content... Nothing useful there??

I'm genuinely curious how does that data help. What would the prompts be like? "Help me design an addictive UX"? How do comments like birthday wishes or people posting their beach pictures and people replying with how good they look add any kind of value to the ML model training? Those conversations would be in larger quantity than any that discuss anything meaningful.

Re: Big Tech's underground race to buy AI training data

#132
post #119

Earlier quoted context omitted.

It's a shiny toy — it'll yield worse answers. Much like Google's own AI.

Google search is terrible. Chatgpt is definitively better for searching right now, and i often find myself reaching for it over google for a wide category of questions.

Google search is terrible because Google's stopped caring about search quality in favor of monetization. It doesn't mean an LLM can outperform a traditional search engine that cares about said quality.

Re: Big Tech's underground race to buy AI training data

#133

>in talks with multiple tech companies to license Photobucket's 13 billion photos and videos >Photobucket declined to identify its prospective buyers, citing commercial confidentiality. >tech companies are also quietly paying for content locked behind paywalls and login screens, giving rise to a hidden trade in everything from chat logs to long forgotten personal photos from faded social media apps In this market, et…

Photobucket is a morally bankrupt shell of its former self. They send constant emails with extremely urgent subject lines threatening to delete your photos unless you sign up for a $5/mo plan. They do this even if your account doesn't contain any photos .

> threatening to delete your photos unless you sign up for a $5/mo plan

What's morally bankrupt about that? It costs money to host your photos and they're a business that can decide to charge their customers any rate they think the market will accept.

Re: Big Tech's underground race to buy AI training data

#134
post #50

Earlier quoted context omitted.

> Google is limited to mainly Android users https://www.appmysite.com/blog/android-vs-ios-mobile-operati... Random link. Can't vouch for it. But US and RoW have quite different patterns.

> Random link. Can't vouch for it Seems about right to me, Android dominates the mobile World by sheer numbers. But what is the value that they can derive from user data? A million Bangladeshi's texts from food delivery is probably a lot less valuable than say a Singaporean using Numbers on Mac OS to layout the next lucrative investment and the data they;d get from the correspondence of say 100 high net worth individ…

> But what is the value that they can derive from user data?

What, are the pictures and videos of people from the global south somehow not good enough to train AI due to their economic situation?

Re: Big Tech's underground race to buy AI training data

#135

Earlier quoted context omitted.

Whatsapp chats are encrypted, how can they be used to train the models? Also what kind of training can be done on Instagram data, is there anything of value there?

> Whatsapp chats are encrypted While they claim E2E encryption, I seriously doubt they would offer this service entirely for free with having some backdoor or potential MITM breach that they likely tucked away in the ToS given the wide use of it it most of the World who pay for SMS/text messages: it just seems so incredibly unlikely to be entirely encrypted from a company that willing gave DMs to Netflix, used Cambri…

> While they claim E2E encryption, I seriously doubt they would offer this service entirely for free with having some backdoor or potential MITM breach that they likely tucked away in the ToS given the wide use of it it most of the World who pay for SMS/text messages: it just seems so incredibly unlikely

You don't have to trust Metas self-regulation, but you best believe the EU does not fuck around on such issues. Self-preservation is a hell of a motivator.

Re: Big Tech's underground race to buy AI training data

#136

Earlier quoted context omitted.

These are probably pretty arbitrarily priced so someone thought they'd be cute & everyone else picked up this rate.

I don't think market prices usually work this way. Am I missing a reason that this is an exception?

Do you really think that market prices are going to reflect such a popular adage because a picture really is worth 1000 words? Does this kind of pricing differential get reflected in the salaries of journalists vs photographers? No? Then as someone else said it's a new market with an artificially "for lolz" initial price that competitors are blindly copying to avoid having to do their own price discovery. I'd expect this to shift & correct slowly over time unless it's a cute joke that everyone appreciates & the true price is close enough that no one cares enough to differentiate in that way.

Re: Big Tech's underground race to buy AI training data

#137
post #106
post #84

Earlier quoted context omitted.

It's not scraping, it's indexing and linking out to creators. LLMs are helping themselves to everything with no regard for content creators. They should be subject to copyright claims — I don't care if it destroys their business, they should've considered that at the outset. They didn't then and they don't care to now, they're simply greedy and looking to build something that benefits themselves and their investors w…

but how can you prove that your picture of a cat was used in LLM? if you owned a franchise called "Chicken Brothers" with a the logo of two chickens standing side by side with arms crossed proudly then do you have claim over all derivatives including the spanish name generated by LLM? i just dont think its straight forward, the main complaint should be payout for license used during training but its tough to prove un…

That's OpenAI's problem and the burden should be on them.

Re: Big Tech's underground race to buy AI training data

#138
post #60

I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.

This won't be necessary in future AIs. As AIs will start aligning tokens from all the rich modalities of audio, video, 3D with text so that they can express complex ideas, they will bootstrap in proper language generation. I don't think college essays, etc would contain anything novel. Future techniques could smoothly interpolate better creating ever-anew wordmud.

I agree with your overall point that an AI which can learn about the world directly won't need eleventy billion documents to learn language generation. Just two comments:

1) Based on how pre-verbal children learn, one nitpick is that I strongly suspect we need to give AI touch and a sense of space in order to truly understand quantity, causality, object permanence, etc.

2) Something that is not a nitpick: even a superhuman multimodal AI wouldn't have direct access to human emotions, sexuality, ideas of natural beauty, etc. I don't think humans have run out of interesting things to say about these ideas.

(In particular, I don't think a superhuman AI is capable of understanding music unless it is directly emulating the biological processes by which humans understand music. The issue is not "logical" - melodies don't actually make sense analytically.)

Re: Big Tech's underground race to buy AI training data

#139
post #60

I wonder if they’ve considered hiring people to write. A lot of people might do it for cheap just to have their imprint on AI. Or another twist pay people to submit ten years of emails (upload the backup file) or just pay small amounts for works they’ve made. College essays, journals, etc.

Turnitin will have millions of essays written by students. No doubt they will already be looking at these deal (or getting ready to update their license if it currently doesn't permit it).

Re: Big Tech's underground race to buy AI training data

#140

Earlier quoted context omitted.

> Random link. Can't vouch for it Seems about right to me, Android dominates the mobile World by sheer numbers. But what is the value that they can derive from user data? A million Bangladeshi's texts from food delivery is probably a lot less valuable than say a Singaporean using Numbers on Mac OS to layout the next lucrative investment and the data they;d get from the correspondence of say 100 high net worth individ…

> But what is the value that they can derive from user data? What, are the pictures and videos of people from the global south somehow not good enough to train AI due to their economic situation?

> What, are the pictures and videos of people from the global south somehow not good enough to train AI due to their economic situation?

I don't make the rules, in fact if you are seriously wondering what use 'darker' people's data have had with AI training look no further than the surveillance based platforms that are responsible for tons of false incarcerations of mainly black US citizens [0].

I'm not sure if it's going to change for the plight of the 'Global South's' data either. It's not that I think it's inherently prejudiced, either; it's more like it's optimized to be greedy in order to extract as much value as it possibly can from the current system at all costs.

People need to stop smoking hopium and thinking that this is going to usher some sort of egalitarian renaissance, this is business as usual by the mega corps that bring you this tech.

0: https://innocenceproject.org/artificial-intelligence-is-putt...

Post reply on HN