atkrista
21 hours ago
I would just LOVE to see all the behind-the-scenes shithousery both companies are employing to one-up the other in this, largely, 2-horse AGI race. Someone should make a mockumentary when all is said and done!
nbardy
21 hours ago
I think it's weirdly just a choice of deciding to cut releases.
We already know OpenAI has "bel" that is MUCH better than astra and is being used internally
233mhz
20 hours ago
> We already know OpenAI has "bel" that is MUCH better than astra and is being used internally
I mean, isn't it almost a guarantee that what we get is a gimped version of what they use internally? They probably already serve themselves next gen level models at 1k+ tps from cerebras machines hosted on perm while we get quantized astra/opus at 50tps on a good day
iLoveOncall
21 hours ago
> We already know OpenAI has "bel" that is MUCH better than astra and is being used internally
You're just believing their own bullshit. There's no indication that this is true except from claims from people working at OpenAI.
If they really had a much more powerful model, it would make absolutely no sense to sit on it.
_davide_
19 hours ago
> it would make absolutely no sense to sit on it.
Yeah, it does, it might be misaligned, a snapshot of the going on training, bigger than they can serve publicly, not yet completed the full training pipeline.
If you see the knowledge cutoff you can see that sol 5.6 finished the initial main (+stage edited) of the training pipeline on Feb 16, 2026 but it was publicly released on July 9.
The opposite would be weird: if they do NOT have an unreleased in-house model that would be really odd.
nextaccountic
16 hours ago
> The opposite would be weird: if they do NOT have an unreleased in-house model that would be really odd.
Indeed it would be really odd if OpenAI were actually open
iLoveOncall
17 hours ago
> The opposite would be weird: if they do NOT have an unreleased in-house model that would be really odd.
This isn't at all what I said. The original commenter mentioned a model MUCH stronger than Astra.
adamzenith
21 hours ago
You don't think having a more intelligent model they can use internally that others can't is an advantage?
owebmaster
21 hours ago
If that's the case, do Anthropic have an even better one helping them? The Chinese labs too? Where this OpenAI "advantage" is taking them?
iLoveOncall
21 hours ago
No? The top of human engineers are much better than any model would be, so AI models really aren't a big advantage when you're trying to develop anything that is SOTA.
x187463
19 hours ago
You can't fathom how 'top human engineers' could take advantage of exclusive access to a frontier model at ultrafast speeds and limitless token budgets to conduct research?
iLoveOncall
13 hours ago
I have all that and LLMs invariably produce garbage, so, no.
ceejayoz
21 hours ago
Not every problem is best addressed by a top engineer.
Plenty of the tasks that keep a company running can benefit from good-enough (and better than the competition).
iLoveOncall
21 hours ago
Yes, and none of those tasks require even the current SOTA models.
233mhz
20 hours ago
> If they really had a much more powerful model, it would make absolutely no sense to sit on it.
Makes total sense if they don't have the compute and can't serve it in an economically viable way. Also it lets you build things no one else in the world can build as fast as you until it's released
owebmaster
5 hours ago
This theory would be easy to prove IF openai was pushing high quality software. It's not.
iLoveOncall
20 hours ago
> Also it lets you build things no one else in the world can build as fast as you until it's released
"We can generate slop faster than anyone in the world" :evil_emoji:
simonw
21 hours ago
It makes sense for them to sit on it until they've finished testing it. More powerful but also more likely to delete all your email by mistake = you shouldn't release it yet.
meowface
21 hours ago
With all due respect, you do not have a single clue what you're talking about.
owebmaster
21 hours ago
With the same respect, you don't either. Simping for openai don't make you part of their in group
meowface
21 hours ago
I definitely do not know what I'm talking about, but that individual doesn't know what they're talking about even more than I don't know what I'm talking about.
howdareme9
21 hours ago
its not done training, why would they release a model that hasn't finished training?
besides, we know anthropic are sitting on models too
iLoveOncall
21 hours ago
> its not done training, why would they release a model that hasn't finished training?
Because clearly they have no problem with releasing newer versions of models even just a week apart.
> besides, we know anthropic are sitting on models too
It is from your crystal ball or from other bullshit you heard from Anthropic employees on Twitter?
We all know Anthropic had Mythos and Fable, and they turned out to be completely normal models, entirely in line with the capability of their predecessors.
All they do is lie, and you're believing their lies.
meowface
21 hours ago
The poster is not claiming Bel is a secret AGI. Just that it exists and only exists internally at the moment.
It's rumored to be over 10T parameters. When released it'll probably be very good at certain tasks, albeit slow and expensive and not necessarily "wiser". You don't have to make this a binary.
Also, Mythos was in fact a significant step-up in several ways. It fits the trend line, but only because the trend line for LLMs is quite steep. Plus Astra is still in many ways less intelligent than Fable/Mythos despite being released much later.
azan_
21 hours ago
> its not done training, why would they release a model that hasn't finished training?
> Because clearly they have no problem with releasing newer versions of models even just a week apart.
You can see how it is pure non-sequitur, right?
> We all know Anthropic had Mythos and Fable, and they turned out to be completely normal models, entirely in line with the capability of their predecessors.
Fable was absolutely not in line with capabilities of other models when released. For cybersec work it was much, much better.
mFixman
21 hours ago
Any strong enough model with weak enough safeguards can cause an AI Chernobyl event that will make people and governments against AI development and deployment, just like Chernobyl did for nuclear energy.
ChromeUltron
21 hours ago
tell me "I drank the kool aid" without telling me you drank the kool aid.
mFixman
21 hours ago
The US government and most large companies drank the kool aid, and they will be the ones blaming Big AI if things go very wrong.
esseph
20 hours ago
I think there's going to be a brutal backlash against the entire technology sector.
Gov and Corp will throw their hands up and explain how it's not their fault.
petesergeant
21 hours ago
> largely 2-horse
The absolute frontier is largely 2-horse, but the rest of the pack is very close behind, which I'm grateful for. Grok, Facebook, and the Chinese vendors are producing excellent models.
bayindirh
21 hours ago
Gemini is also pretty nice for researching things. It turns out that having the whole internet indexed and having unlimited access to YouTube is a force multiplier of some kind.
Since Google has their own TPUs, TPS is also pretty high w.r.t. Claude, for example.
thunfischtoast
21 hours ago
I think Gemini could be a great product, if they didn't decide to stuff it down my throat at every possible occasion.
Recent example: on my e-reader, tapping a word I don't now and clicking "Translate" pulls up the possible translations from a dictionary, a local file just a couple of megabytes big, near instantly, on this tiny processor.
Doing the same on my Android phone starts a Gemini-chat with the prompt "Translate the word x into y". Takes forever, internet access needed, results vary, burns who knows who much energy.
Why? Just why?
selestify
21 hours ago
So that some team at Google can hit their OKRs for user adoption and get promoted.
bayindirh
21 hours ago
That kind of shoving down is pretty bad, I agree.
I neither use Android devices nor Google Search, so the only Gemini thing I see is the Gemini chat interface.
I understand the pain, though.
lp92
19 hours ago
Since Gemini 3.8 Flash release, I've been largely using that for my coding ttasks. The architectureand design work I use Opus, but the rest of the work is done via Gemini 3.8. I don'thave to worry aboutweekly limits even on the $20 plan. I get down to about 10% weekly limit before it resets. Currently I'm building my own generative modeling + aero sim package largely using Gemini and it's been working great.
lxgr
21 hours ago
Gemini is unbelievably bad for research in my experience. It hallucinates like it's 2023, doesn't use its own search, makes up fake rationalizations for why it didn't need to etc.
It's baffling that OpenAI managed to get better at web search than the company literally synonymous with web search.
bayindirh
21 hours ago
Pretty interesting. It one-shots correct information with references and a further reading list 99% of the time for me.
I don't want it to fill in the gaps though, but make it reference anything and everything it brings, hence it doesn't hallucinate much.
If something feels off, I ask it to back it with concrete data, and if it can't, I don't consider that information correct. That happened once, though, and web doesn't have any information on that thing either. So in that case, not only it had no information on the web, the training data had no information on that thing either short of feeding confidential design documents if they were ever present in the first place.
The question was about an instrument preamp though, so nothing crucial.
lxgr
19 hours ago
I guess there's one thing that could really bias my anecdata: I almost exclusively use Gemini in a professional/niche domain and ChatGPT in a personal/idle curiosity context, so there's a chance that ChatGPT is just as wrong but I am not knowledgeable in a given domain to tell.
Bluestein
21 hours ago
... and, must be said a plethora of largely unsung, small, unknown "labs", outfits, "researchers" and the like. There is a long tail of smart people having at this. I guess, sheer compute aside, I think much progress - or, at least, important pieces thereof, will come from there.-
bayindirh
21 hours ago
There are some niche research areas where bog standard machine learning algorithms make miracles. LLM is just the poster child. AI/ML is a much larger and wider research area.
Bluestein
21 hours ago
Your "bog" if unrelated, put me in mind of a "Cambrian explosion" (of intelligence) in this case, ongoing.-
In a way: We have had "recursive self improvement" (RSI) - that's what genetics is, giving rise to ... us.-
This time around, I think the truly worrying thing is that, having transcended the biological substrate, the pace is unlike anything previously seen.-
Just a thought.-
michelb
21 hours ago
I REALLY need new seasons of 'Silicon Valley'
tom1337
20 hours ago
Still waiting for the moment where GPT either orders 4,000 pounds of raw beef or just deletes the whole OpenAI repo because it "thought the easiest way to remove all bugs is by deleting the whole repository"
TeMPOraL
21 hours ago
[flagged]