ks2048
9 hours ago
> It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results
> Basically the paper is so horribly written that it’s impossible to read it without AI help
That's interesting and haven't seen this in all the coverage of this event.
It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.
TheOtherHobbes
9 hours ago
Math proofs need to produce the correct output correctly, which is not quite the same thing.
This looks like an AI IPO PR powerplay, because at this point the proofs haven't been checked and it may not be possible for a human to check them - because proofs should be clear, not horribly written and noisy.
The noise is suspicious because it's the difference between brute forcing and cognition. A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.
You want the path through the maze to be as short as possible and the map to be as clear as possible.
This sounds like the opposite. There may be a genuine path through the maze, but if it's too convoluted and takes too long it will be impossible to confirm.
I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
I suspect that's possible without tripping over the halting problem. (But I can't prove it.)
eadler
8 hours ago
That reminds me of this paper:
Chow, T. Y. (2008). A beginner’s guide to forcing (arXiv:0712.1320). arXiv. https://doi.org/10.48550/arXiv.0712.1320
> “All mathematicians are familiar with the concept of an open research problem. I propose the less familiar concept of an open exposition problem. Solving an open exposition problem means explaining a mathematical subject in a way that renders it totally perspicuous. Every step should be motivated and clear; ideally, students should feel that they could have arrived at the results themselves. The proofs should be “natural” in Donald Newman’s sense [13]:
> This term . . . is introduced to mean not having any ad hoc constructions or brilliancies. A “natural” proof, then, is one which proves itself, one available to the “common mathematician in the streets.””
kens
7 hours ago
> I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
In 1976, the proof of the Four Color Theorem was controversial because it was done with a computer examining over 1000 cases by brute force and was essentially not comprehensible by humans. But mathematicians ended up accepting it. So mathematics has a 50-year precedent of not requiring human-scale proofs. How is the current situation different?
(Disclaimer: Apologies if this sounds dismissive or argumentative. I genuinely think that the Four Color Theorem should play a role in these discussions and suspect that many people are unaware of the controversy over it.)
pizza234
7 hours ago
There's also another (that I find more concerning) aspect to it.
As AIs become smarter and smarter, there will be no amount of clarity that will make more complex proofs understandable to humans - this is an inevitable effect of the cognitive capacity gap.
Complaining about bad style can make some sense now (I disagree anyway), but it's an argument that will be dead shortly.
dist-epoch
6 hours ago
Arguably no human understands 100% of how a smartphone is produced, and, it doesn't matter?
Maybe no human will fully understand a future proof, but they could fully understand a little piece of it. And many humans in aggregate could understand it, each with their own little piece.
FloorEgg
8 hours ago
If intelligence is compression, and these models are a different form of lesser intelligence than human, but being scaled up to brute force problems, then it makes sense the artifacts that produce (the proofs) would have worse compression than a human proof would.
In other domains I have seen first hand overwhelming evidence of how things that cause the AI to make mistakes also cause humans to make the same mistakes.
I wonder if the proofs being produced that are hard for humans to interpret are also hard for other LLMs to interpret.
In other words, I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.
It kind of an explicit example of how the LLMs can be materially less intelligent than people, but still be more productive through scaling, and yet they also can't replace people because they are a categorically different kind of intelligence. It's like all the AI debates compressed into one example showing countwr-intuitive answers.
perching_aix
6 hours ago
> I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.
Isn't it fairly established that (generally [0]) manually written / optimized skill files perform a lot better than generated ones? Meaning that yes, this likely does hold.
[0] or to be specific, that the pecking order is: ai generated < human co/written < hyperoptimized for the specific model via some convergence process
esafak
8 hours ago
This is just the first cut. I have no doubt that they will polish their proofs over time.
slopinthebag
8 hours ago
idk if i'd even say they're "lesser", just very different. so they look like gods/babies depending on what they're doing because we anthropomorphise them.
FloorEgg
6 hours ago
In terms of synapse density, neuron diversity, energy efficiency, memory access, etc. they are orders of magnitude lesser.
My mental model - for better or worse - is that intelligence has both shape and area, and LLMs are orders of magnitude smaller area but very different shape, and they have more intelligence area in the kind that humans have lesser of.
So yes very different, more in some material ways and lesser in others, but in total intelligence are still orders of magnitude lesser.
My gp comment was acknowledging that when you scale up many instances / brute force problems it confuses that "total area" claim a bit.
To follow the anthropomorphization... 1000 toddlers may have more total intelligence than a grown man, but does that matter?
The problem with these discussions probably/usually fold into differing/loose definitions of intelligence.
slopinthebag
6 hours ago
ok yeah then i think we're in total agreement
perhaps the chat-based ux has sort of fooled us into comparing these things to human intellegence. we don't really do this with chess, or other forms of ai, nor computers at large.
Octoth0rpe
8 hours ago
> A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.
https://en.wikipedia.org/wiki/Inter-universal_Teichmüller_th... seems like a counterpoint, but IANAM. (I am likely cherrypicking the far end of the bell curve re: straightforward here)
ffaccount2
7 hours ago
Not a counterpoint, actually case in point, because:
>Mochizuki and a few other mathematicians claim that the theory indeed yields such a proof but this has so far not been accepted by the mathematical community.
Proof can't be understood, proof doesn't matter.
IsTom
7 hours ago
Isn't this controversial, to say the least?
dist-epoch
6 hours ago
I wait for AI to say something about this :)
Someone at OpenAI, please, work on this.
pizza234
8 hours ago
The post says there's a Lean certificate for this and other proofs ("some [...] not all of them").
> This looks like an AI IPO PR powerplay,
Interestingly, the post has actually also an argument for this:
> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
> So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.”
smcg
8 hours ago
It's on OpenAI and Anthropic to prove that they obtained these results legitimately and credited all researchers who deserve credit. They do not get the benefit of the doubt.
Kotlopou
8 hours ago
But if you think they got them illegitimately, then how did they get them? And why are mathematicians reacting to this as a sudden explosion of new results that have resisted sustained effort? Where is the sudden productivity rise coming from?
za_creature
7 hours ago
I'd say it comes from the same mathematicians that were strongly encouraged to use the machine to solve their problems for the last 2 years or so.
There's clear benefit in a babelfish that can coordinate disparate efforts, the only problem with the current iteration is giving credit to said efforts.
Google went quite far down the road to hell, but stopped short of taking credit for websites' content since the company understood that poisoning the well only goes so far. At this point, one can safely conclude that _Chat_GPT was an intentional attempt to squeeze out more data once they mined the internet dry.
pizza234
7 hours ago
Have you actually read the article? It's been actually written, among the other things, because the author's wife has been trying to solve one of the problems for her whole life.
sebzim4500
8 hours ago
Surely by the time of the IPO we will know whether the main results are correct, if only because a different AI will have produced a lean proof or found a logical flaw (the second case would be hard to verify but probably not impossible).
Also from what I can tell from the few fields I understand, the proofs aren't that long or complicated they are just terribly written.
curt15
8 hours ago
Why should that make material difference to the IPO? What is the economic value of those results?
The entire US federal budget for math research is something like $100M annually. And mathematicians in other countries are hardly making bank either. How does one reconcile how the market has historically valued mathematics with the cash-strapped frontier labs ploughing so much money into that enterprise?
coderenegade
6 hours ago
They're burying any doubt that the models are capable of superhuman performance on intellectual tasks. Neural nets aren't calculators, and were notably poor at mathematical reasoning tasks for a long time. Now they're not, and the labs are proving that by chewing through what would ordinarily be decades of progress in a month. And the reason to go for math in particular is because there's no wiggle room. You can't just dismiss it as hallucination.
If the models can do this, they're almost certainly good at just about everything, because the reasoning and creativity required to solve these problems will translate. And even if they were only good at this stuff, that's still a tremendously valuable thing, because quantitative reasoning and analysis is the bedrock for many, many industries.
oAI is gunning for the largest IPO in history at this point, and they might actually get there.
curt15
5 hours ago
> If the models can do this, they're almost certainly good at just about everything, because the reasoning and creativity required to solve these problems will translate.
The "then" in your "if-then" bears a heavy load. Why would society assign so little economic value to pure mathematics if the skills for proving math theorems translate to massive value in "just about everything"? Would you expect top mathematicians to cure cancer if you transplanted them from the math department to a medical research lab?
coderenegade
17 minutes ago
Because research mathematics is a subset of all quantitative work that gets done, but it's by far the most technically difficult subset. If the models can handle research grade math and produce ironclad proofs, they can probably handle the quantitative side of just about any discipline in a trustworthy fashion. Think about how many dinky spreadsheets have gone on to become critical tooling for large organizations. Even if you consider that many disciplines hide technically demanding work behind tooling (e.g. essentially no one is writing a stiffness matrix FE routine by hand), this model would be capable of writing a direct competitor from scratch to produce the same result.
All of STEM relies on mathematical analysis, and new models are now superhuman at that. And yeah, I'd go a step further and say that the reasoning and creativity required to solve cutting edge math problems probably does translate to other tasks like interpretation of the law, or medical diagnosis, or accounting, etc., for the same reasons that I think most top tier mathematicians would excel at those tasks were they so inclined.
ndriscoll
3 hours ago
Do you think mathematicians are not already working on cancer research? There's quite a bit of heavy math in biostatistics, medical imaging, machine learning, etc.
I'd assume that the majority of people who study math take their skills and move onto some related STEM career that isn't pure math. Academia is incredibly small and competitive.
runarberg
8 hours ago
The market works in mysterious ways. What companies do for marketing is often irrational, what companies do to attract investors is likewise often irrational, and why investors invest in companies is also often irrational.
Why should that make a material difference to the IPO? Because of the vibes, and investors are indeed all about the vibes.
jryle70
3 hours ago
It works in mysterious ways, but you know exactly it will behave certain way "Because of the vibes, and investors are indeed all about the vibes."?
ComplexSystems
8 hours ago
> I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
Who do we demand this from? The AI companies? Or the mathematicians who are worried they will have nothing left to do?
za_creature
8 hours ago
From the entity that is producing these proofs, obviously.
As the old saying: great claims require great evidence.
esafak
8 hours ago
Ask away. They've dropped the mic, as far as they're concerned; they're not going to worry about what you do with it, or if you don't understand it.
za_creature
7 hours ago
That's the best definition of slop I've ever read.
caaqil
8 hours ago
We should consider the possibility that at some abstraction levels, we can safely stop chasing "clarity" or "coherence" which is circularly defined in such a way that it's capped by human processing power.
Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running. Mathematicians will get there.
TheOtherHobbes
4 hours ago
I've very aware that human cognition has limits.
But it doesn't follow that these proofs - or any proofs - are automatically on the far side of that limit.
The human usefulness of a proof depends entirely on its human legibility. Much of the value of proofs is in inventing new techniques and concepts and having new insights into relationships. Occasionally you get some game changing insight into practical physics or engineering. But that's rare.
Without that, proving or disproving a conjecture is an excuse for new and original thinking.
Compilers are not the same problem. The point of code is to produce reliable-ish consequences from various possible inputs. It's not a creative exercise in logical consistency, which is what maths proofs are, ultimately.
jltsiren
8 hours ago
CS got that idea from mathematics. Theorems (with the definitions required to state them) are supposed to be self-contained units. Once the general consensus is that a theorem has been proven correct, people can use it without understanding the proof. Of course, people still want to understand how things work, and it often makes sense to understand them a couple of layers below the one you usually work at. But at some point, you should stop distracting yourself with irrelevant details and focus on your actual work.
yorwba
7 hours ago
Thing is that many of the theorems here are not useful work in and of themselves, but were posed as research problems because it wasn't clear how they could be resolved with current techniques, implying that the process of trying to find a proof might result in new techniques. It's those new techniques that are the actual goal, but if they can't be easily extracted because the proof isn't structured to enable this, that's a bit of a headache.
ssfdg
8 hours ago
This proof dump reminds me of the glut of low-quality drive-by PRs overwhelming open-source repos.
dormento
8 hours ago
Its like infinite summer of code, but for math. Must be annoying.
piker
9 hours ago
It also aligns with the fear that these proofs present a risk to the ecosystem by out-competing attempts at more human-readable proofs. Perhaps though we end up with more math influencers who edit and annotate these proofs to bring them back to us.
whatshisface
9 hours ago
The ecosystem is (ahem) gated by hiring committees. There is no risk of AI replacement from the inside. "Replacement" is not even a possible movement. The funding for mathematics worldwide comes mostly from endowments, which are investment pools.
bobajeff
9 hours ago
I think that's ultimately a good thing. As proofs weren't supposed to be the point as stated by William Thurston long ago. Maybe now the focus can be more on better explanations and creating tools for growing understanding and intuition.
cowlevel
8 hours ago
Good explanations should take the form of human-understandable proofs.
btilly
8 hours ago
Define "human understandable".
It's worthy of note that most humans, do not find most mathematicians understandable. As is frequently demonstrated in Calculus classes. Therefore it is arguable that even human produced results are not generally human understandable.
rrr_oh_man
9 hours ago
Vibe mathing
amazingman
3 hours ago
Sounds a lot like using LLMs for coding ~18mo ago.
rramach
7 hours ago
True but the explanations will get better.
Meanwhile you can use the model to help you out as Scott comments "Just now, however, Dana tells me that she’s been asking Astra all day to explain the new proof of the UGC to her and it’s been doing an amazing job and she’s starting to understand the construction."
jltsiren
8 hours ago
Isn't that just the default experience with AI these days? In small enough scale, AI models can express their ideas clearly. But the larger and more complex the ideas are, the less suitable the outputs are for human consumption. I guess AI models think too different from humans, and nobody has trained them to communicate complex ideas in the way human experts in that particular topic expect.
KoolKat23
7 hours ago
This makes sense and looks exactly what something would look like that is smarter than us. If it is correct it is making inferences that we can't see. At a stretch even working in more dimensions than our three dimension limited brains.
acedTrex
8 hours ago
> Basically the paper is so horribly written that it’s impossible to read it without AI help
This basically describes every single PR at work for the past year. Diffs of 10k+ paragraphs of comments saying nothing. Just rubber stamp and move on, nothing else you can do.
spelunker
8 hours ago
I see many parallels to genAI-assisted code development. Not surprising I think.
smrtinsert
7 hours ago
Sounds like exactly the daily I deal with when agents try to create product requirements from multiple sources, except to a much much worse extent.
m3kw9
8 hours ago
why not get Astra to make it make sense?
aaroninsf
9 hours ago
Serious question:
Why would anyone believe this (also) is not simply example N+1 of this is the worst it will ever be, as opposed to recognizing this as what will almost certainly prove to be an awkward moment, soon to be replaced by another order of magnitude of cleaner, clearer, more intelligible, etc.?
Ximm's Law: every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon.
devin
9 hours ago
Devin's Law: every defense of AI which rests on "it will get better, trust me" is in many ways indistinguishable from 2010s crypto hype or "level 5 self driving is right around the corner"
usrnm
8 hours ago
1) Predicting the future is hard, but so far everyone who was saying that it would get better turned out to be right. It is getting better 2) Waymo exists
devin
7 hours ago
3) That doesn't mean flying cars will within your lifetime
I don't think anyone is saying it can't or won't get better, but the question is how much better, on what timescale, and are there fundamental parts of the problem which will remain extraordinarily difficult to improve?
The comment I was responding to suggested a guarantee of an "order of magnitude" jump right around the corner. There is no guarantee of this, and if you view doomers as fools for having doubts, then we ought to look upon the folks who are sure of this sort of progress in the same way.
kulahan
7 hours ago
Your original comment was about self-driving cars, not flying ones. If you wanted unattainable goalposts you should’ve started with that, not ended with it.
devin
7 hours ago
Waymo does not claim level 5 self-driving, and you don't have a level 4 in your driveway, so there was no moving of goalposts.
kulahan
7 hours ago
Then say that, instead of bringing up flying cars, which is moving them, explicitly.
devin
6 hours ago
Respectfully, I disagree with the way you're characterizing my comment. I was demonstrating that just because you have a level 4 Waymo doesn't mean flying cars are right around the corner. This is again a reference to the original comment I was replying to, the one that suggested of course we're going to get an order of magnitude improvement.
Kotlopou
8 hours ago
In that case, one would expect to see some progress in this direction, but AFAICT that hasn't shown up yet? If anything, it's getting worse, though that could just be the increasing scale and decreasing cleanup efforts.
Already the unit distance proof was substantially human-edited (per Thomas Bloom). Then with the ten problems from Astra you started getting the citation issues. Then Navier-Stokes was a rushed 160 pages with barely any citations, and some of the related papers were called (by their "authors") the ugliest mess they've ever seen.
And now here we are. At least it seems that mathematical ability and communication with a mathematical audience are independent skills, and progress in the first does not imply the second.
This doesn't surprise me much, given two analogies: 1) many smart people are nonetheless horrible lecturers. (You can't quite get the opposite extreme, since to explain math well you have to be able to do it.) 2) AI writing in general hasn't improved. The models have annoying verbal tics ("honestly") and have no sense of which part of what they say is obvious and which is relevant.
auggierose
8 hours ago
You can get the opposite extreme quite often as well, I'd think. How many really good lecturers have never proven a new important result?
Kotlopou
7 hours ago
I'm thinking of somebody like Grant Sanderson (3blue1brown), doing pure exposition extremely well. For that you at least need to be able to work through examples, or to present why an intuitive approach might fail, and these things can be little theorems themselves. It doesn't have to be publishable in the current culture of novel results, but you do need a lot of competence with the tools.
auggierose
4 hours ago
Yes, Grant Sanderson might be a great example. I am not doubting the competence of the "good lecturers" I was referring to. But that is different from being able to introduce the big new concepts that give the big new results.