Grok is a surprisingly good automated theorem prover

1 pointsposted 6 hours ago
by henryrobbins00

Item id: 49010310

1 Comments

henryrobbins00

5 hours ago

Some additional details about these results that didn't make it into the main post:

- The OpenATP "standard provers" were used; see docs [5] for model / harness configuration details

- Time and cost are function of effort level, which may lead to unfair comparison across provers

- FATE-X excludes task 10 since claude and grok hit session limits

- FATE-X excludes leanstral and aristotle due to temporary endpoint failures

- Deepseek's FATE-X accuracy is corrected from 2 to 3 due to verifier bug (now fixed)

- 2 FATE-X deepseek misses are sorry-free, but rely on native_decide

- Claude's FATE-X miss is due to a failed delegation to a background subagent

- All costs come from underlying CLI, except codex which uses pricing table