That was my initial reaction, however, there is more to it.
The thing that made these breakthroughs feel important was that they seemed to be open ended problems that the system itself made independent progress on. That the results were not primarily great retrieval into the corpus of mathematical research + stitching together.
A major point of having novel frontier problems as a benchmark itself is to try to sidestep this issue a bit, where it's hard to tell if one is evaluating reasoning vs retrieval capabilities.
But at least one of the problems was identified as definitely being plausibly mostly retrieval, since it hinges on but does not cite a 2016 paper that the author identified. Closely enough that the author describes it as plagiarism. Another problem seems to combine results from 2016 and 2019.
Ignoring the question of credit, it makes it look like OpenAI doesn't actually understand the eval well in the first place, and undercuts the notion that the results represent great leaps forward in mathematical reasoning.
I mean are they wrong to want credit? This isn’t like “I brought coffee to the meetups and corrected some spelling errors” but more like OAI pulling the “you made this? I made this” meme.
You can dig through every single paper and find some citation that was missed. No one cares that much unless the result is important, and then you get bitter recriminations saying that it was stolen or plagiarized. Tale as old as time. See e.g. Schmidhuber, who has a long list of vendettas against people he thinks stole his work.