Every model (incl. Jev) we tested inflates security finding severity

7 pointsposted 4 hours ago
by brene

3 Comments

brene

4 hours ago

Eleanor from my team and I collab'ed on this benchmark. A few things worth stating up front because they shape how to read this. We constantly see LLMs saying "THIS IS A CRITICAL SECURITY FINDING" and we wanted to put it to the test. Most LLMs naturally gravitate towards using CVSS for severity scoring.

We wanted a task where model judgment could be checked against verified results from security engineers (ground truth) and root cause "why" security findings are always inflated. TL;DR LLMs still make too many severity judgements without the proper context, so naturally they bias towards the "worst case" scenario. Happy to answer any questions on this topic

user

2 hours ago

[deleted]