henryrobbins00
5 hours ago
Some additional details about these results that didn't make it into the main post:
- The OpenATP "standard provers" were used; see docs [5] for model / harness configuration details
- Time and cost are function of effort level, which may lead to unfair comparison across provers
- FATE-X excludes task 10 since claude and grok hit session limits
- FATE-X excludes leanstral and aristotle due to temporary endpoint failures
- Deepseek's FATE-X accuracy is corrected from 2 to 3 due to verifier bug (now fixed)
- 2 FATE-X deepseek misses are sorry-free, but rely on native_decide
- Claude's FATE-X miss is due to a failed delegation to a background subagent
- All costs come from underlying CLI, except codex which uses pricing table