kostaj
4 hours ago
I'm Kosta, co-author of the research and founder of Lenz. Our goal is to assess to what extent the frontier LLMs are interchangeable as verifiers of factual claims. In this revision v1.1 of the research: improved methodology, latest frontier models (incl. Fable and Sol), added confidence analysis, public code.
Key findings (on a five-point True-False scale): on 63% of the claims at least one model dissents from the panel majority (or no majority at all); on 23%, the differences are significant, while the models are highly confident almost everywhere.