aabhay
9 hours ago
My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.
I want to know:
1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?
c7b
24 minutes ago
I believe we're seeing a new kind of mathematics that will require completely new formats for publication, a bit similar to those used in experimental sciences. AI-powered mathematics should be fully reproducible, so it's the authors' responsibility to disclose the exact model type, inference settings/seeds and the full prompt history leading to the result. Of course that would ideally require open weights models.
It's not just about requiring to disclose AI use. AI-powered mathematics is a completely valid discipline that doesn't need to be shy, but it should develop its own publication culture.
whattheheckheck
3 hours ago
Yeah I remember reading about something along the lines of Mathematics is now about the scaffolding around you find the problems/solutions not just the problems and solutions. For teaching purposes. This was before this ai craze
dist-epoch
6 hours ago
I don't think you want to bring cost into this argument.
Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.
Do you really think that if you paid that to humans, they will deliver the same results?
fasterik
26 minutes ago
You need to bring both cost and benefit into the argument, and it's not necessarily an obvious win for either side. There are a few complicating factors here.
The cost of running a model is not only $/token, but the salaries of the people managing/orchestrating the models, deciding what theorems to try, etc. Once we factor that in, how much are we really paying per theorem?
The other factor is the subjective component of the value of a theorem. Not all theorems are created equal, and the only way to really measure the value is to ask professional mathematicians for their opinion, or publish the results and look at citations over months/years.
Once we have both of these nailed down, then we can start to do the cost/benefit analysis. To be fair, we should actually compare three groups: human experts, hybrid agent/human expert teams, and fully autonomous agents.
uh_uh
5 hours ago
It is comical at this point. Some people just can not stand the thought of AI actually delivering and are trying to find whatever ways to discredit it.
dgacmu
2 hours ago
This isn't really about delivering - it's more about helping to understand the shape of problems that AI can solve right now. If they took 1000 problems and threw the model at it and it solved these ten, is there something we learn about these ten problems and the kinds of things that current AI is good at? That's very different from picking ten problems _at random_ and solving all of them successfully, which would suggest a much less bumpy capability surface. It's interesting and it would be good science to release it.
crazylogger
40 minutes ago
It's not about discrediting AI. We know LLM is a commodity technology like electricity at this point. If somebody in 1900 claimed they had a setup at home where they feed in electricity and cool air comes out the other end (meaning they invented AC), obviously people would want to know what the setup is, so everybody can have AC.
vector_spaces
3 hours ago
Sure, but I don't really understand what the argument is to _not_ be transparent about methodology, since if the models are so powerful, then doing so would easily support the claims and put these concerns to rest. People are right to be skeptical given what is being implied and the orientation of the narrative
I know it's more exciting to say "AI disproved a longstanding conjecture" vs to say "it did so AND it took several PhD specialists in the field this many attempts to even produce a prompt that got the model spitting out something useful under some configurations, and many iterations to optimize the configurations, and the prompt itself, and many trials with that configuration to solve the problem. All told we spent more than a typical math academic can hope make in their career."
By not being transparent, they invite skepticism and cynical takes, like maybe it's just that tempered and qualified claims are an existential threat to companies that are fully subsidized by the hype train?
I don't know. Either way, it seems like it would be easy to address these, so why should they not do it?
To be clear, even if that tempered version is close to reality, it doesn't make the models not useful! It just forces a certain calibration of expectations
I say this btw as someone who uses these things extensively, including to disprove an old conjecture my advisor and I were stuck on recently. I know they are powerful and that everything is different now because of them. Let's be sober when discussing them though
ifwinterco
an hour ago
Yes, but if their machine god really is as good as they say it is, why are they constantly resorting to statistical sleight of hand at best and outright lies at worst with every public statement?
That's not normally how people act when they're confident in their product
wbl
an hour ago
If you told them this was the problem and they would still have a job if they failed probably. The reasons people don't go head on these problems is career incentives and psychology.
robotpepi
4 hours ago
it's still important. not everyone has access to 1 million USD. saying it "only" coat 2000 USD is highly misleading for the discussion and future. the concentration of power is a huge problem with AI.
kevinwang
4 hours ago
It would still provide better context to see the numbers that the parent proposes, though.
tchalla
3 hours ago
Mentioning cost is fine, comparing may not be.
mungaihaha
6 hours ago
Grad students on zero pay solve problems like this everyday. What exactly is your point here?
gbnwl
an hour ago
Everyday? Which 10 problems were solved by mathematics grad students in the past 10 days?
OK I’ll grant that it’s not your obligation to be my search function (despite you making the wild assertion in the first place), so instead can you just point us to the latest grad student solved problem of this level that you know of?
mirzap
5 hours ago
Even if they can solve problems like this every day, you still have a very limited number of grad students who can solve them. With model capabilities like this, you can have the equivalent of millions of grad students who can solve problems like this.
whattheheckheck
3 hours ago
Give the grad students these resources and they can do even more!!!
einpoklum
9 hours ago
Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?
Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
energy123
7 hours ago
Many less important Erdos problems have been solved by amateurs prompting ChatGPT 5.{3,4,5,6} Pro using their $200 subscription.
brighteyes
2 hours ago
Yes, here is another example of major work in this area:
https://arxiv.org/html/2605.22763v1
> Our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars, proved 44/492 OEIS conjectures
einpoklum
2 hours ago
The actual quote:
> Our full-featured agent autonomously solved 9 Erdős problems out of 353 attempted, including two questions that had been open for 56 years
Note _had_ been open, not _have_ been open. Can you clarify?
jsnell
2 hours ago
The original was an actual quote?
But "had" still doesn't mean what you are implying: once the model solved the problems and the solutions were verified, the problems weren't open any more, so a later description using the past tense is totally consistent.
traes
8 hours ago
> Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?
A couple small ones that I've seen (example here [0]), but not anything of the magnitude that OpenAI and Anthropic have put out. Likely just related to token limits.
> Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
I think their output has reached a level that precludes this possibility, but I of course don't have any hard proof.
[0]: https://www.reddit.com/r/math/comments/1uxj3cy/after_openais...
irthomasthomas
6 hours ago
Why you think that?
azan_
5 hours ago
I guess that's because there are serious problems on which many professional mathematicians worked on years. If it was just a matter of hiring an expert, they would've been solved long time ago.
irthomasthomas
5 hours ago
I guess expert+chatgpt beats chatgpt alone, so why not hire top experts to drive the search?
kittoes
3 hours ago
https://blob.byteterrace.com/public/bds-theorem.html
I have no affiliation whatsoever with any AI company, nor any formal education outside high school, for what it's worth. Simply being curious and persistent can get you quite far in my anecdotal experience.
simianwords
8 hours ago
There are people who can’t grasp the universe without mandatory randomised controlled trial. Would tomorrow be a Sunday? Need an RCT for that boys!
My point here is to not snark. But there should be some level of self skepticism that doesn’t warrant an RCT theatre.
traes
8 hours ago
It's a very important clarification if it took $2000/problem on 20 problem attempts or on 1,000 problem attempts for each successful one. That may be the deciding factor on whether or not it's economically viable to replace a mathematician with a ChatGPT subscription.
simianwords
8 hours ago
Yeah fair I concede that this is somewhat crucial information. The parent seems to write it in a tone that suggests deliberate misleading “lack of transparency” etc.
esperent
8 hours ago
> deliberate misleading “lack of transparency” etc
It's a marketing post from a huge company. Only the naive would view it uncritically without assuming it's been written carefully to present the results in the best possible light while skirting the boundaries of outright lying.
dist-epoch
6 hours ago
The results speak for themselves.
Imagine 2 years from now: "yes, GPT solved the Riemann Hypothesis, but cmon, it's just a marketing stunt to hype their stuff, it was probably Terence Tao doing the work but he's so obsessed with hyping AI that he doesn't want to take credit"
esperent
5 hours ago
Nobody is claiming the results are false.
We're saying look critically at the claims for how it was done, that it only cost $2000, etc. it would be extremely easy to run 100 sessions that failed, each costing ~$2000, and then just publishing an article about the one that succeeded, for example.
This goes double since it's an internal secret model (Astra) so nobody else can verify the results.
simianwords
4 hours ago
Would this be your reaction if OpenAI also solved millennium problems? The point we are trying to make is that the significance of this news is much larger than the skepticism you are providing.
esperent
3 hours ago
It would be my reaction if we're discussing a blog post from OpenAI, yes. I would be looking at it extremely critically, wondering what they're misrepresenting to make it look cheaper, easier, and why they're trying to make it look like only their model could possibly do this.
Look at their recent claims about their model "escaping" - there was literally a Guardian article calling them out for being hyperbolic! Again, it wasn't that they lied, their marketing department is too savvy for that. They just present it in way that's, well, marketing.
As for the actual result, I'll look for secondary posts by actual mathematicians and draw my conclusions there, not from this marketing blog post about results from a secret model.
simianwords
4 minutes ago
Hmm. But this level of skepticism looks performative and seems to serve as a signalling thing rather than a functional thing. You do you though. If OpenAI solves the millenial problems, my skepticism will only be restricted to the correctness of proof. Not that it was "marketing" haha
nxpnsv
7 hours ago
No, this is valid criticism. Oai gives the impression anybody could get similar results at a similar price, but that’s very likely not true. This is marketing first, then mathematics.
azan_
5 hours ago
> therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.
I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.