> In particular, Astra’s key mathematical step combined ideas first found in two papers from 2016 and 2019. Andreas Thom, a mathematician at the Dresden University of Technology, who co-authored both papers, summarized the result on MathOverflow.com, calling it “creative and at the same time elementary.”
Looking at the mathoverflow post, he says:
```
I think it is too early to judge if this construction will serve any other purpose than providing the right framework to apply these earlier results efficiently. On the other side, I was looking myself for such a mechanism ever since we wrote the paper in 2019 and admire the efficiency of this construction.
```
With machine automated formalizations now feasible, I wonder whether git repos (which would automatically include attribution) wouldn't make more sense as modern journals. e.g. you could have the core mathlib as the "prestige" journal, and then various "results in X" repos as more niche journals? If a theory or result becomes central enough, it is elevated to the core. Peer review becomes a PR, etc. "Journals" could also run a timestamp authority/notary that records when a commit was pushed.
This would make it much easier to catch such things and also easier to build more powerful math agent harnesses. Like you could imagine something like a Lean Hoogle that can locate any applicable theorems to whatever type(s) you have as a tool call.
That was my initial reaction, however, there is more to it.
The thing that made these breakthroughs feel important was that they seemed to be open ended problems that the system itself made independent progress on. That the results were not primarily great retrieval into the corpus of mathematical research + stitching together.
A major point of having novel frontier problems as a benchmark itself is to try to sidestep this issue a bit, where it's hard to tell if one is evaluating reasoning vs retrieval capabilities.
But at least one of the problems was identified as definitely being plausibly mostly retrieval, since it hinges on but does not cite a 2016 paper that the author identified. Closely enough that the author describes it as plagiarism. Another problem seems to combine results from 2016 and 2019.
Ignoring the question of credit, it makes it look like OpenAI doesn't actually understand the eval well in the first place, and undercuts the notion that the results represent great leaps forward in mathematical reasoning.
I mean are they wrong to want credit? This isn’t like “I brought coffee to the meetups and corrected some spelling errors” but more like OAI pulling the “you made this? I made this” meme.
You can dig through every single paper and find some citation that was missed. No one cares that much unless the result is important, and then you get bitter recriminations saying that it was stolen or plagiarized. Tale as old as time. See e.g. Schmidhuber, who has a long list of vendettas against people he thinks stole his work.
It’s extremely aggravating that AI firms are given the easy task of selling and advertising a product that has obvious potential and uses but feel the need to over hype more and more to the point where the actual uses aren’t even the focus anymore but instead the focus is on what the product (and the firms) fail to do.
I know that part of that is scrambling to find a way to justify the insane debt and margins they need to make up for but I feel that it could have been done.
> In particular, Astra’s key mathematical step combined ideas first found in two papers from 2016 and 2019. Andreas Thom, a mathematician at the Dresden University of Technology, who co-authored both papers, summarized the result on MathOverflow.com, calling it “creative and at the same time elementary.”
Looking at the mathoverflow post, he says:
``` I think it is too early to judge if this construction will serve any other purpose than providing the right framework to apply these earlier results efficiently. On the other side, I was looking myself for such a mechanism ever since we wrote the paper in 2019 and admire the efficiency of this construction. ```
So it maybe is a pretty interesting result.
With machine automated formalizations now feasible, I wonder whether git repos (which would automatically include attribution) wouldn't make more sense as modern journals. e.g. you could have the core mathlib as the "prestige" journal, and then various "results in X" repos as more niche journals? If a theory or result becomes central enough, it is elevated to the core. Peer review becomes a PR, etc. "Journals" could also run a timestamp authority/notary that records when a commit was pushed.
This would make it much easier to catch such things and also easier to build more powerful math agent harnesses. Like you could imagine something like a Lean Hoogle that can locate any applicable theorems to whatever type(s) you have as a tool call.
Nothing more real in research than academics fighting over credit and missed citations.
That was my initial reaction, however, there is more to it.
The thing that made these breakthroughs feel important was that they seemed to be open ended problems that the system itself made independent progress on. That the results were not primarily great retrieval into the corpus of mathematical research + stitching together.
A major point of having novel frontier problems as a benchmark itself is to try to sidestep this issue a bit, where it's hard to tell if one is evaluating reasoning vs retrieval capabilities.
But at least one of the problems was identified as definitely being plausibly mostly retrieval, since it hinges on but does not cite a 2016 paper that the author identified. Closely enough that the author describes it as plagiarism. Another problem seems to combine results from 2016 and 2019.
Ignoring the question of credit, it makes it look like OpenAI doesn't actually understand the eval well in the first place, and undercuts the notion that the results represent great leaps forward in mathematical reasoning.
I mean are they wrong to want credit? This isn’t like “I brought coffee to the meetups and corrected some spelling errors” but more like OAI pulling the “you made this? I made this” meme.
You can dig through every single paper and find some citation that was missed. No one cares that much unless the result is important, and then you get bitter recriminations saying that it was stolen or plagiarized. Tale as old as time. See e.g. Schmidhuber, who has a long list of vendettas against people he thinks stole his work.
It’s extremely aggravating that AI firms are given the easy task of selling and advertising a product that has obvious potential and uses but feel the need to over hype more and more to the point where the actual uses aren’t even the focus anymore but instead the focus is on what the product (and the firms) fail to do.
I know that part of that is scrambling to find a way to justify the insane debt and margins they need to make up for but I feel that it could have been done.
Paywalled.
Is it? Seems to work for me?