prasadvara
5 hours ago
This is great writeup, can we do a cross comparison with "human" experts?? whether models perform better OR worse??
LambdaComplex
4 hours ago
This writeup reads like it was written by Claude, which makes me immediately question its accuracy.
> Each dot is one code sample; bars mark the median. The split between malicious and benign packages, perfect before the prune, was perfect after.
People do not write like this.