Live data from Hacker News

4TB of voice samples just stolen from 40k AI contractors at Mercor

app.oravys.com

31–40 of 250 posts

Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor

#31

Earlier quoted context omitted.

Data hoarding predates LLMs. There where other machine learning methods which also needed data for training.

“Before LLM’s there was_____” I see this whenever an LLM’s impact is assessed. We know. The issue is scale and the ability for smaller and smaller groups (down to individuals) to execute at scale. Fake news always existed. Now one dude in India can flood multiple sock puppet media accounts with right wing content/images (actual example) at a scale previously unimaginable.

I really hate this when it's something negative that humans also do. It's like, yeah, people do do that, but why are we automating {negativeTrait}?

Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor

#32

Earlier quoted context omitted.

“Before LLM’s there was_____” I see this whenever an LLM’s impact is assessed. We know. The issue is scale and the ability for smaller and smaller groups (down to individuals) to execute at scale. Fake news always existed. Now one dude in India can flood multiple sock puppet media accounts with right wing content/images (actual example) at a scale previously unimaginable.

Do LLMs require that much more data than the tradional ML approaches we've seen over the years?

Yes. This is pretty well established. Neural networks in general are considerably less sample-efficient than traditional ML methods. The reason they became so successful is that they scale better as you increase training data and model size. But only with modern compute power they became useful outside of academic toy model applications.

Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor

#34

I wonder how many of the current text-to-speech ML models have large parts of leaked or "stolen" data in their training data? Almost none of the TTS releases seem to talk about exactly where they get their training data from, for some reason. I also wonder if we'll see an explosion in SOTA TTS in ~6 months from now.

Not really, Mozilla Common Voice (the ImageNet of speech) is larger than this. Their English database has 3814 hours, 1.6 million sentences, from 100k speakers.

https://commonvoice.mozilla.org/en/languages

Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor

#35
post #3

The only data that cannot be stolen or leaked is data that doesn't exist. Hard lesson for both users and companies. Germans (because of course) have a word for this: "Datensparsamkeit". Being frugal with your data.

Data that is publicly available also can't be stolen or leaked. Nobody can steal Mozilla's common voice dataset.

Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor

#36

Man that’s pretty shitty that Mercor tricked 40k contractors, and then did a poor job of securing their data. There should be stronger consequences for stuff like this.

What happens now is that a lot of clueless CTO that didn't know about this company now know it's name. So the outcome of this mess is probably more business for Mercor

I mean, just look at what happened to Crowdstrike....

Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor

#37

Earlier quoted context omitted.

I miss the pre-LLM days when you could make a decent argument that having any unnecessary data was just a liability. Now all anybody thinks is “more data for the AI!”

Were you not around for the Big Data heyday a decade ago?

Hell you mean a decade ago? I still see businesses running losses left right and center saying that they're gonna monetize user data, any day now.

Related "monetizing user data" seems to just mean ads. Ads on everything, forever, until the userbase gets fed up and moves to a new service that definitely won't do that, and the cycle repeats about every 3 years.

Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor

#38
post #3

The only data that cannot be stolen or leaked is data that doesn't exist. Hard lesson for both users and companies. Germans (because of course) have a word for this: "Datensparsamkeit". Being frugal with your data.

Data can never be stolen, because it is not a physical thing. Data can be copied, and it can be erased - sometimes both happens at the same time. Data can be lost, that is when its last existing copy was erased.

Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor

#39

Earlier quoted context omitted.

Were you not around for the Big Data heyday a decade ago?

Until thumb drives became large enough to fit most datasets it stopped becoming Big Data. Just normal data.

We have thumb drives that can store petabytes of data?

Or did you mean the "big data" crowd which thought 500GB was noteworthy? I don't think anyone took those serious, neither in 2010s nor now. That was always "small" data

Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor

#40
post #3

The only data that cannot be stolen or leaked is data that doesn't exist. Hard lesson for both users and companies. Germans (because of course) have a word for this: "Datensparsamkeit". Being frugal with your data.

> Germans (because of course)

I don't know if it's the reason you imply. In the 70s, there were big debates in Germany about privacy and data storage. They spoke of one's data shadow (Datenschatten). I suspect this word comes from that tradition. The reason the word exists would then be the reflection (Verwaltigung) on WW2.

Post reply on HN