Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets

34 pointsposted 13 hours ago
by 882542F3884314B

13 Comments

croemer

10 hours ago

In principle interesting, but I can't stand the Claude writing.

> Keys with real blast radius

> Here is what they unlock.

> This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used.

> We cloned the public dataset hub end to end: every repository, every branch, every large-file object

> The size is only half the story. These are the training sets behind models people actually use. The worst-hit ones are named, card-documented pretraining corpora that open models were built on. We verified every credential we cite against its provider, so they were live when we looked.

The whole post looks like a Claude artifact with random little cards.

It's also just too long, which is a side effect of using LLMs, it's just too easy to create walls of text.

fractorial

9 hours ago

I similarly find it interesting; I do not understand why it is not the default to just generate the thing as a draft, research anything you aren’t clear on, and re-write it in your own voice.

xena

9 hours ago

That takes a level of taste and craft that many in this industry do not possess.

manquer

6 hours ago

It takes a level of taste and craft to write good content.

It only bit of effort to write in your own voice , anyone with high school level of writing skills should be able to muster if they were inclined to do so .

gumby

9 hours ago

Ever read a book by Yuval Noah Harari? He seems to have been writing like this since before LLMs. Perhaps that’s where they got it from.

jgalt212

8 hours ago

reply with ASD-STE100 -> Yuval

ks2048

8 hours ago

I think "7.6 PB" is more informative than "4.4x Empire State Building heights worth of DVDs", but that's just me.

lorreyfum

10 hours ago

Wouldn’t it just be easier to crawl the net? Not sure what huggingface has to do with anything here.

croemer

10 hours ago

I guess Huggingface datasets are easy to crawl - those datasets are hosted to be crawled. In contrast to the web at large which will be behind Cloudflare etc

Izmaki

8 hours ago

This is fine. All of this is fine. :’)

user

10 hours ago

[deleted]