BREAKING: OpenAI might have stolen another major proof.
In a detailed Mastodon post, which I report in full in the comments, Andreas Thom presents several pieces of evidence suggesting that OpenAI may have trained Astra on conversations in which he and Gábor Kun were working on Gromov’s soficity conjecture, one of the ten problems OpenAI later announced Astra had solved.
I know Andreas. We met several times early in our careers. He is an exceptional mathematician, a leading expert on sofic and hyperlinear groups, and one of the most respected scholars in the field. He has spent two decades working on this problem.
If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself.
And if the allegations raised by Levent Alpöge, Tristan Buckmaster, and now Andreas Thom are all substantiated, we are no longer looking at isolated incidents.
We may be looking at one of the greatest intellectual scandals in the history of science.
AI is not discovering new mathematics.
AI is stealing human discovery.
Not Navier-Stokes, that was a couple days ago, old news. It’s one of the less famous problems they posted around the same time.
OpenAI is guilty of computer crime through massive unauthorized access to web servers to scrape their content, evading every kind of blocking attempt and crashing servers all over the place. This is the same thing Aaron Swartz was prosecuted for, but nothing seems to be happening to OpenAI. The content being public is irrelevant to this crime. You could have a public domain book in your bedroom, but if I break into your house and copy it, I’m a burglar even if I’m not a copyright infringer.
Meta (Facebook) is known to have done a huge copyright infringement by downloading a massive pirate library (libgen) to train its AI. That’s separate from the claim that using the texts for AI training is infringing in its own right (there are lawsuits about that going on, and it’s not a slam dunk issue imho). I remember this reported specifically about Meta but it would shock me if OpenAI and Anthropic didn’t do the same thing. So they have no business whining about model distillation.
Separately from that, yes, the stuff in the math prompts was supposed to be private but OpenAI apparently trained on them anyway. OpenAI is a multi-tasker and can do more than one bad thing at the same time.
Well, one obvious example is licensing. You can have things that are presented online with a terms of service or an actual license that excludes certain uses.
One specific example would be GitHub and the sites like it. There are plenty of public facing code bases with licensing that would prohibit forms of reuse, like for financial gain.
And yes, this appears to be what the researchers believed to be private conversations, but are likely completely owned by the AI companies in their terms of service.
How does one steal what is public?
This article is about stuff that was private, yes?
OpenAI is guilty of computer crime through massive unauthorized access to web servers to scrape their content, evading every kind of blocking attempt and crashing servers all over the place. This is the same thing Aaron Swartz was prosecuted for, but nothing seems to be happening to OpenAI. The content being public is irrelevant to this crime. You could have a public domain book in your bedroom, but if I break into your house and copy it, I’m a burglar even if I’m not a copyright infringer.
Meta (Facebook) is known to have done a huge copyright infringement by downloading a massive pirate library (libgen) to train its AI. That’s separate from the claim that using the texts for AI training is infringing in its own right (there are lawsuits about that going on, and it’s not a slam dunk issue imho). I remember this reported specifically about Meta but it would shock me if OpenAI and Anthropic didn’t do the same thing. So they have no business whining about model distillation.
Separately from that, yes, the stuff in the math prompts was supposed to be private but OpenAI apparently trained on them anyway. OpenAI is a multi-tasker and can do more than one bad thing at the same time.
Well, one obvious example is licensing. You can have things that are presented online with a terms of service or an actual license that excludes certain uses.
One specific example would be GitHub and the sites like it. There are plenty of public facing code bases with licensing that would prohibit forms of reuse, like for financial gain.
And yes, this appears to be what the researchers believed to be private conversations, but are likely completely owned by the AI companies in their terms of service.