GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index

21 pointsposted 9 hours ago
by wertyk

12 Comments

NiekvdMaas

9 hours ago

Title: "major gains"

First chart: from score 61 (GPT-5.6 Sol) to drumroll 61 (GPT-6 Astra)

flyaway123

9 hours ago

Indeed. Though to be fair it is referring to "Artificial Analysis Coding Agent Index", from 65 to 67.

dist-epoch

8 hours ago

I think they mean cost per task, where Astra is now on the Pareto frontier.

aogaili

3 hours ago

I don't understand how those labs are releasing models so close in performance to one another?

Are they just scaling more? getting more data at the same rate? training against the same benchmarks? making the same breakthroughs?

How can this be explained?

bashtoni

6 hours ago

Are the Artificial Analysis benchmarks really worthwhile any more?

They really don't seem to match my real-world experience, and based on the comments I see I don't think that match most other people's either.

For example, Opus 5 was at the top for some time. My experience is that it's not noticeably better than Opus 4.8, and it definitely seems worse than Fable 5, which AA benchmarks put behind Opus 5. GPT 5.6-sol and Opus 5 seem pretty interchangeable, although Sol is noticeably better at finding problems in code, particularly edge cases.

villish

an hour ago

I have no faith in these benchmarks. Muse 1.3 shouldn’t even be in the same conversation yet it scores above GPT 5.6 & 6.0

eis

8 hours ago

In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58). In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3.

Am I missing something or is this not looking too... stellar?