Earlier quoted context omitted.
Data hoarding predates LLMs. There where other machine learning methods which also needed data for training.
“Before LLM’s there was_____” I see this whenever an LLM’s impact is assessed. We know. The issue is scale and the ability for smaller and smaller groups (down to individuals) to execute at scale. Fake news always existed. Now one dude in India can flood multiple sock puppet media accounts with right wing content/images (actual example) at a scale previously unimaginable.
4TB of voice samples just stolen from 40k AI contractors at Mercor
31–40 of 250 posts
Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor
#32Earlier quoted context omitted.
“Before LLM’s there was_____” I see this whenever an LLM’s impact is assessed. We know. The issue is scale and the ability for smaller and smaller groups (down to individuals) to execute at scale. Fake news always existed. Now one dude in India can flood multiple sock puppet media accounts with right wing content/images (actual example) at a scale previously unimaginable.
Do LLMs require that much more data than the tradional ML approaches we've seen over the years?
Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor
#33Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor
#34I wonder how many of the current text-to-speech ML models have large parts of leaked or "stolen" data in their training data? Almost none of the TTS releases seem to talk about exactly where they get their training data from, for some reason. I also wonder if we'll see an explosion in SOTA TTS in ~6 months from now.
Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor
#35The only data that cannot be stolen or leaked is data that doesn't exist. Hard lesson for both users and companies. Germans (because of course) have a word for this: "Datensparsamkeit". Being frugal with your data.
Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor
#36Man that’s pretty shitty that Mercor tricked 40k contractors, and then did a poor job of securing their data. There should be stronger consequences for stuff like this.
I mean, just look at what happened to Crowdstrike....
Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor
#37Earlier quoted context omitted.
I miss the pre-LLM days when you could make a decent argument that having any unnecessary data was just a liability. Now all anybody thinks is “more data for the AI!”
Were you not around for the Big Data heyday a decade ago?
Related "monetizing user data" seems to just mean ads. Ads on everything, forever, until the userbase gets fed up and moves to a new service that definitely won't do that, and the cycle repeats about every 3 years.
Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor
#38The only data that cannot be stolen or leaked is data that doesn't exist. Hard lesson for both users and companies. Germans (because of course) have a word for this: "Datensparsamkeit". Being frugal with your data.
Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor
#39Earlier quoted context omitted.
Were you not around for the Big Data heyday a decade ago?
Until thumb drives became large enough to fit most datasets it stopped becoming Big Data. Just normal data.
Or did you mean the "big data" crowd which thought 500GB was noteworthy? I don't think anyone took those serious, neither in 2010s nor now. That was always "small" data
Re: 4TB of voice samples just stolen from 40k AI contractors at Mercor
#40The only data that cannot be stolen or leaked is data that doesn't exist. Hard lesson for both users and companies. Germans (because of course) have a word for this: "Datensparsamkeit". Being frugal with your data.
I don't know if it's the reason you imply. In the 70s, there were big debates in Germany about privacy and data storage. They spoke of one's data shadow (Datenschatten). I suspect this word comes from that tradition. The reason the word exists would then be the reflection (Verwaltigung) on WW2.