Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
We scanned every public dataset on Hugging Face, covering 7.6 petabytes and 187 million files. Our analysis uncovered over 221,000 live credentials, including keys for cloud infrastructure, databases, and AI providers like OpenAI and Anthropic. These leaks expose sensitive data and pose significant supply chain risks, with potential annual losses exceeding $920,000 from unauthorized AI inference usage.
A single top-tier key drained to its cap bills more in one month than this entire 7.6-petabyte scan cost to run.
- croemer
In principle interesting, but I can't stand the Claude writing.
> Keys with real blast radius
> Here is what they unlock.
> This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used.
> We cloned the public dataset hub end to end: every repository, every branch, every large-file object
> The size is only half the story. These are the training sets behind models people actually use. The worst-hit ones are named, card-documented pretraining corpora that open models were built on. We verified every credential we cite against its provider, so they were live when we looked.
The whole post looks like a Claude artifact with random little cards.
It's also just too long, which is a side effect of using LLMs, it's just too easy to create walls of text.
- ks2048
I think "7.6 PB" is more informative than "4.4x Empire State Building heights worth of DVDs", but that's just me.
- lorreyfum
Wouldn’t it just be easier to crawl the net? Not sure what huggingface has to do with anything here.