Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
trufflesecurity.com
Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
1–10 of 14 posts
Re: Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
#2> Keys with real blast radius
> Here is what they unlock.
> This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used.
> We cloned the public dataset hub end to end: every repository, every branch, every large-file object
> The size is only half the story. These are the training sets behind models people actually use. The worst-hit ones are named, card-documented pretraining corpora that open models were built on. We verified every credential we cite against its provider, so they were live when we looked.
The whole post looks like a Claude artifact with random little cards.
It's also just too long, which is a side effect of using LLMs, it's just too easy to create walls of text.
Re: Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
#3Re: Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
#4Re: Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
#5Wouldn’t it just be easier to crawl the net? Not sure what huggingface has to do with anything here.
Re: Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
#6In principle interesting, but I can't stand the Claude writing. > Keys with real blast radius > Here is what they unlock. > This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used. > We cloned the public dataset hub end to end: every repository, every branch, every large-file object > The size is only half the story. These are the training sets behind models pe…
Re: Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
#7In principle interesting, but I can't stand the Claude writing. > Keys with real blast radius > Here is what they unlock. > This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used. > We cloned the public dataset hub end to end: every repository, every branch, every large-file object > The size is only half the story. These are the training sets behind models pe…
I similarly find it interesting; I do not understand why it is not the default to just generate the thing as a draft, research anything you aren’t clear on, and re-write it in your own voice.
Re: Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
#8In principle interesting, but I can't stand the Claude writing. > Keys with real blast radius > Here is what they unlock. > This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used. > We cloned the public dataset hub end to end: every repository, every branch, every large-file object > The size is only half the story. These are the training sets behind models pe…