Something Is Wrong at OpenAI

3 pointsposted 7 hours ago
by vincent_s

7 Comments

verdverm

7 hours ago

I was thinking more about the cultural problems that lead to stories like this

https://globalnews.ca/news/11676795/tumbler-ridge-school-sho...

Perhaps what OP is noticing is that models are plateauing for real world use cases (even while doing even more impressive things when you have unlimited tokens to burn), and their training regime to try anything and everything until "the task is complete" to the point of throwing spaghetti at the wall to see what sticks

vincent_s

7 hours ago

I think they try pushing the models in a direction where they just go on until somehing is finished end-to-end. So basically building "the loop" into the model itself. E.g. with GPT-5.5 the model would often do only part of a job and get back to me with a recommendation for next steps and I would either say "Yes, go on" or "No, instead do X". And the newer models never even get back to me, they just keep on churning, deciding about the direction on their own. Sometimes they go into the right direction, sometimes they drift into something unrelated or just go down some rabbit holes for hours. I liked GPT-5.5 better, and I built my own tools around it to still get them to finish tasks end-to-end which works fine.

vincent_s

7 hours ago

Btw. I absolutely agree with your analysis that models seem to be more or less plateauing for real world use cases and that OpenAI is pushing them into the "try anything and everything until "the task is complete"" direction. They should have stopped at GPT-5.5 and rather tried to make this one cheaper and faster. The rest is then the job of surrounding tools and/or the "harness". For my SaaS [0] I am still using GPT-5.5 with great success and have no plans of upgrading to GPT-6.1/Sol. GPT-5.5 is good enough and gets the job done reliably. GPT-6.1/Sol is fine as a reviewer of you take it's output with a grain of salt and not follow it verbatim.

[0] https://vroni.com/

verdverm

6 hours ago

I'm not completely convinced we are technically plateauing for real world problems, but they are being trained to optimize for other things (IPOs), and in this we get the impressive results (eg. math proofs) and "hacking"

verdverm

7 hours ago

this is why open weights is the way

1. run it yourself, token vendors cannot discontinue it

2. fine-tune for personal preference

3. just as good for almost everything these days

vincent_s

7 hours ago

Absolutely. But "run it yourself" is completely underestimating how hard it is to run inference at scale for large models and at the same time bringing costs down. The experience (speed, costs, cache) I have with various providers and models on OpenRouter is way worse than using Codex and Claude and if you want the outcome to be comparable in quality it also costs a lot more than using OpenAI's and Anthropic's subscriptions.

verdverm

6 hours ago

that's open router, not the way to buy open weight tokens, setting aside the 5% surcharge

use OpenCode Go or Fireworks instead