rsfern
4 hours ago
Regardless of what you think of the priority dispute issue discussed on sibling threads, I’m highly skeptical of the closing quote that this Navier Stokes result means that the same approach of casually spending a few million on agentic computation is going to solve end to end materials design or drug development.
Those problems can’t be formally verified with an automated theorem prover. We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations because otherwise they’d be too computationally expensive, or we just don’t have the right data to parameterize them beyond describing qualitative behavior. Agents are helping accelerate research in these fields but I think it’s mostly a different class of problem that’s a lot harder to specify and verify
rakejake
4 hours ago
Yeah, I think you can't just throw money randomly at problems and expect results unless you know a line of attack that can get you all the way. OpenAI chose the line of attack only after it became known to them via rumors. They "front-ran" the researchers.
eieje1
10 minutes ago
You can’t do anything novel with these models from scratch and let it fly. I’ve observed something over the past few months
Work on something novel -> llm is kinda useless and low value-add -> Keep at it and in the process feed it more information -> keep doing this periodically -> a few months go by and you realise the model outputs are almost like-for-like regurgitations of what was inputted in some prior period.
Once it’s accumulated new info can it produce something automated that is somewhat useful? Sure.
But by itself - absolutely not.
I clearly see humans will be needed - the best ones that is. For ‘rote work’ and stuff that is not IP sensitive firms will be ok with employees putting that as inputs into models.
But I’d wary about trusting the labs. They will push the letter of the law to the max.
Personally I’ve stopped doing anything novel with these models. If I do use a model on something adjacent but not totally novel I have to craft the inputs in a strategic way not to give much away.
I’d wager firms will soon realise this and that growth rate of revenues of the frontier labs will become questionable.
cmiles8
2 hours ago
Yes. What the headlines hailed as an AGI discovery the facts show more to be someone spending years mining for gold, rumor gets to OpenAI that there might be gold in this specific place, they mine there and instantly discover gold, then tell the world they’ve developed the worlds best gold finding/mining machine.
Separate from all the allegations of more nefarious actions and ethical issues, that’s the most charitable version of what happened here.
Romario77
22 minutes ago
they threw it on all the millenial math problems (I think there are 6 at this point unsolved, well, 5 now).
And according to them at some point they saw that one was close to being solved, so they pointed all the agents at it.
The same thing happens to humans - at this time there are no simple problems left, so solving the hard ones requires using prior knowledge and attempts at solving things.
dcre
3 hours ago
Worth noting they claim they did not choose the line of attack. Of course we don’t know whether that is true.
rakejake
2 hours ago
Plausible deniability - The line of attack is in their sessions/prompts data. Just make the prompt pointed enough that the search space is tractable and use your ginormous compute.
> "Of course we don’t know whether that is true"
Yep. Who is verifying these claims? We all know how trustworthy Altman & Co are.
whimsicalism
an hour ago
but the researchers were also largely relying on AI
cmiles8
an hour ago
“relying on” is misleading here relative to what the researchers have said.
If I write a book and pass it through a spelling and polish checker, I still wrote the book and its core IP. I didn’t “rely on” the tool to create the IP.
whimsicalism
an hour ago
it’s much more like you come up with the premise and someone else writes the book. the released prompts for other foundational problems (like unit distance) prove that.
vouaobrasil
34 minutes ago
The tools the researchers used though was much more than an spellchecker, because spellcheckers don't come up with chains of reasoning for the arguments in the book. The LLMs did in the case of the Navier-Stokes problem.
pphysch
an hour ago
In the same way you rely on a keyboard or touchscreen to type this comment. It doesn't mean the tool is the brain behind the work.
Romario77
19 minutes ago
that's not how AI was used in this case. It's more like a professor with assistants.
Professor says the assistants - why don't you dig in this direction, I have a hunch it might produce something valuable. And AI assistant does just that, proving or disproving a hunch. This would take the professor a lot of time if doing by themselves.
vouaobrasil
26 minutes ago
Keyboards don't suggest chains of reasoning or words to type. When I press the K key, I know exactly what will happen. It's just a translation layer that gives an output known ahead of time and thus does not impinge upon the creativity of putting words together.
A better example would be playing chess against a player slightly stronger than me and using a chess computer to suggest some good moves. I could win, but it certianly wouldn't be just my brain that wins. It would be an amalgamation of my brain with a machine that suggests good moves.
One cannot simply reason by analogy.
whimsicalism
an hour ago
frankly don’t know how to reply to these sorts of comments anymore
pphysch
an hour ago
That's usually a good sign you are on shaky ground!
logancbrown
30 minutes ago
Obvious false analogy in your earlier argument.
u1hcw9nx
3 hours ago
For any practical application, numerical solvers for Navier-Stokes already exist and do a good job.
This proof is just checking the boxes for mathematicians.
robotpepi
3 hours ago
you're as sure of what you say as wrong about it.
Toutouxc
2 hours ago
Note that your reply has exactly 0 value for anyone who doesn’t already know where and how the parent poster is wrong.
hyperbovine
2 hours ago
The same could be said of your post.
OpenAI (claim to) show the existence of *a* finite time singularity. It could stimulate more research in PDE solving, and maybe physics, but it has zero impact on practical applications, that I can see. The Millenium problems were chosen based on hardness not practical relevance.
CyberDildonics
2 hours ago
If that were true you could explain it. There are lots of solvers for navier stokes simulations and they do a good job.
jgalt212
2 hours ago
which part is wrong?
> For any practical application, numerical solvers for Navier-Stokes already exist and do a good job.
or
> This proof is just checking the boxes for mathematicians.
rsfern
3 hours ago
Agreed, but i think this underscores my point. We have numerical simulations in materials science too, but that doesn’t mean formally verified theorems about the underlying equations automatically translate to formal (or even informal) verification of simulation results. That’s not to say you can’t make progress with agents, but I think it’s less well defined how you write the goal and progress assessment for an agent
harhargange
4 hours ago
I’m pretty sure that OpenAI has some of the best mathematicians prompting the models and analysing the results. While they are marketing as if the model solves problems themselves.
nayroclade
3 hours ago
Prompting them yes, suggesting potentially fruitful research directions and so on, but the actual research was conducted by hundreds of agents swapping millions of messages and using billions of output tokens over 88 hours. The result being a huge Lean proof: https://github.com/openai/NavierStokesAndEuler. It's not just possible for humans to manually guide such a process in a meaningful way. They can set the direction and attempt to understand the result, but they solution itself must emerge (or not) from the agent swarm.
So yes, the models do seem to be "solving" the problems themselves, but not necessarily in the way we think of mathematical discoveries happening. Academic mathematics has historically been resource constrained: There are a limited number of top-level mathematicians, and they only have so much time and brain power to spend. So when approaching a problem, they are essentially forced to be as efficient as possible, not just searching for a solution, but for one that can be achieved within their cognitive budget. This induces them to develop novel techniques and abstractions, and it is actually those techniques and abstractions that tend to be the valuable part for further research, not the proof itself.
An agentic swarm is like getting a single skilled mathematician, cloning them a hundred times, then locking them in a room with the single objective of solving a problem. No longer constrained by time or brain power, they can approach it differently, using pre-existing techniques to gradually build their way to a solution. This process might not require a single intuitive leap or new discovery, and the solution will not be simple or elegant, but they will probably get there. It is more like a process of intelligently guided search than invention.
jhrmnn
3 hours ago
Working with AI on science (not LLMs though), couldn't agree more.
sigmar
3 hours ago
>We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations
Do you think it is possible that better math will lead to better physics models?
tantalor
3 hours ago
It might but the math results from GenAI so far have been limited to finding counterexamples to known conjectures, not building new mathematics.
rsfern
3 hours ago
Yes, definitely! There’s a long history of this and I think there’s tons of opportunities for more. Both for improving the exactness/physical fidelity of models and for developing new approximate theories and simulation methods
alansaber
an hour ago
The TL;DR is still "AI helpful, but not end of the line". The live discussion about these matters is always ridiculously inflated by hyperbole.
jgalt212
2 hours ago
> Those problems can’t be formally verified with an automated theorem prover.
It certainly seems like any problem that is amenable to reinforcement learning will be solved.
rsfern
2 hours ago
It does, yes. So designing objections functions and making sure you can afford the training rollouts becomes really important in defining which problems are tractable. It will be really interesting to see how that shapes the kinds of problems people choose to work on