bob1029
23 minutes ago
> In Hard-Decisions, our benchmark of decision models on multi-step logic
"model": "gpt-6-luna",
"reasoning_effort": "none",
This article seems to be missing important points regarding how these models are intended to be used. It is my understanding that the Decisions API is designed for quick, single-step logic. We already have a proper Death Star for dispatching the more complex problems.I am currently using the Responses API with my clients, which is mandatory to get at non-zero reasoning effort in the latest models. Luna without reasoning turned on might as well be a model from early 2025. This is not how anyone is using this. Responses with 5.6-luna+ and high+ reasoning level feels pretty close to the Star Trek computer experience for me.
Attempting to recreate the OAI reasoning model capabilities at home seems like a pointless quest now. You will never get the access into the base models that the frontier companies have internally. You will also never have access to an engineering team with that kind of capacity. You must submit to the black box if you want the advertised performance figures.