Live data from Hacker News

Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets

trufflesecurity.com

11–14 of 14 posts

Re: Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets

#11
post #8
post #2

In principle interesting, but I can't stand the Claude writing. > Keys with real blast radius > Here is what they unlock. > This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used. > We cloned the public dataset hub end to end: every repository, every branch, every large-file object > The size is only half the story. These are the training sets behind models pe…

Ever read a book by Yuval Noah Harari? He seems to have been writing like this since before LLMs. Perhaps that’s where they got it from.

reply with ASD-STE100 -> Yuval

Re: Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets

#12
post #2

In principle interesting, but I can't stand the Claude writing. > Keys with real blast radius > Here is what they unlock. > This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used. > We cloned the public dataset hub end to end: every repository, every branch, every large-file object > The size is only half the story. These are the training sets behind models pe…

[dead]

Re: Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets

#13
post #7

Earlier quoted context omitted.

I similarly find it interesting; I do not understand why it is not the default to just generate the thing as a draft, research anything you aren’t clear on, and re-write it in your own voice.

That takes a level of taste and craft that many in this industry do not possess.

It takes a level of taste and craft to write good content.

It only bit of effort to write in your own voice , anyone with high school level of writing skills should be able to muster if they were inclined to do so .

Post reply on HN