Show HN: Agent Memory Leaderboard – first public results for AI memory systems

3 pointsposted 10 hours ago
by IreneAI

1 Comments

claudiusa

4 hours ago

Are the published numbers single-run or averaged, and which model does the judging? With LLM-as-judge scoring I would expect a couple of points of run-to-run noise, which does not matter for the top spot but matters a lot for the middle of the table.