OpenAI has a LOT of work to do if they think Luna can compete with Jev

16 pointsposted 11 hours ago
by AnthusAI

6 Comments

bob1029

23 minutes ago

> In Hard-Decisions, our benchmark of decision models on multi-step logic

  "model": "gpt-6-luna",
  "reasoning_effort": "none",
This article seems to be missing important points regarding how these models are intended to be used. It is my understanding that the Decisions API is designed for quick, single-step logic. We already have a proper Death Star for dispatching the more complex problems.

I am currently using the Responses API with my clients, which is mandatory to get at non-zero reasoning effort in the latest models. Luna without reasoning turned on might as well be a model from early 2025. This is not how anyone is using this. Responses with 5.6-luna+ and high+ reasoning level feels pretty close to the Star Trek computer experience for me.

Attempting to recreate the OAI reasoning model capabilities at home seems like a pointless quest now. You will never get the access into the base models that the frontier companies have internally. You will also never have access to an engineering team with that kind of capacity. You must submit to the black box if you want the advertised performance figures.

rileymat2

7 hours ago

Alan is young, round, and kind, but that doesn't mean he isn't also rough and cold at times, as well. … Young round people who are green are usually blue. … Kind people with rough skin are usually red because it's wind burn. If someone shows that they are red, then they are also showing that they are green. …

Statement: Alan is not blue.

A log-probability of −0.00182 is a probability of 99.82%. We asked for five alternatives and got none: Luna put essentially nothing on "true" or "false". And it's wrong. Alan is kind with rough skin, so he's red; red means green; young, round and green means blue. "Alan is not blue" is false, three steps in.

—————-

Can someone explain this I got unknown as well. The problem statement includes the word “usually” a few times.

red369

4 hours ago

I'm interested how this works too.

Is there some sort of specific meaning or rule in this domain that makes this problem mean something different to how it would be read at face value?

Without knowing anything extra, to me this reads:

1. Alan is young, round and kind

2. Alan is sometimes rough

3. Kind with rough skin are usually red

(Alan has not been stated to be in this category - unless "sometimes rough" implies "rough skin")

4. If red, then green

(As above, no information yet on whether Alan is red, so this gives no additional information about Alan)

5. Young, round and green are usually blue

(No information yet on whether Alan is green, so this gives no additional information about Alan)

6. Statement: Alan is not blue = ??

(No additional information since statements 1 & 2: Alan is young, round and kind, Alan is sometimes rough)

BTW - I have just numbered the statements in the order I used them, in case anyone wants to correct or discuss anything. This isn't the order they were given.

Edit: I am forgetting my predicate logic, and didn't recognise this. I think this example has more decoration (is more loosely worded) than I was ever used to. I now think "sometimes rough" and "rough skin" are intended to be interpreted as meaning the same thing.

That makes everything I wrote above this edit wrong. With the statement that Alan is rough, it is implied that Alan is usually blue

itg

7 hours ago

If I'm reading this right, they didn't actually use the Decisions API, they used GPT-6 Luna. Wouldn't call this a good comparison.

sixhobbits

3 hours ago

Is decisions api even rolled out yet?