Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

621–630 of 862 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#622
post #604

Earlier quoted context omitted.

Sure, they have clusters. We often call them "closet clusters". Nothing has changed. Some institutions have larger systems (some extremely large) but none of them have demonstrated running warehouse-scale systems. I'm talking one to two orders of magnitude (and the storage and networking to make sure all those systems don't stall waiting for data).

Perhaps I am underestimating the size of these frontier models then, how big of a cluster do you think would be required to serve a university?

I don't think you're taking this thread seriously enough to respond in detail. A university can be tiny, or it can have thousands of researchers. Some of them might want to work on small models, and others on huge frontier models. Multiple groups training their own models at the same time. Combined with the storage, networking, power, and redundancy, you basically need large data center scale. And at that point you're basically throwing a lot of capital trying to compete with the hyperscalers; you might be able to serve a small number of researchers very well, but most of the consumers would end up unhappy.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#623
post #359
post #354

Earlier quoted context omitted.

you massively collapsed what AI companies have been doing by comparing it to old internet-scraping. Facebook flat-out admitted that they scanned copyrighted books for their AI. The image generators most definitely trained on copyrighted images.

LAION and Common Crawl both scraped copyrighted images. From what I can tell (I'm not an expert in this domain at all), the main difference between those two and frontier labs is in how they stored and used the data. CC and LAION seem to be actually open (unlike "Open"AI) and are more centered around publicly sharing the data they scrape to support research and innovation. OpenAI et al also stole everything from ever…

Common Crawl is text-only.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#624

Earlier quoted context omitted.

what reason do we have to believe that they did this? both things were proved by AI, isn't it logical that they could have very similar approaches? it is common that multiple people essentially simultaneously prove/invent the same thing I see zero evidence of wrongdoing

OAI started working on this only after they found out it was close to being solved. They threw a team of researchers who spent sleepless nights + a ton of compute. This is not exactly healthy academic competition - it's like if you spend a year hunting for oil fields and finally find a very promising area to be explored, only to find that Exxon tapped their entire exploration unit to go all in and and find it overnig…

that is not related to the accusation that they literally stole Buckmaster's work, which seems baseless

I don't agree with the oil claim analogy. this is knowledge, freely given to the world. not something hoarded by a corporation

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#625
post #571

From what I understand, none of the people involved here are originators of the idea that led to this solution. Not Buckmaster nor Alpöge nor OpenAI. All of the above were using LLMs to push other mathematicians' ideas forward (Diego Cordoba and Luis Martinez-Zoroa; named in the linked document). I don't know how the math community handles this but normally I would think if X mathematician comes up with an idea and Y…

You're right that the case is not so clear cut since both sides were using LLMs. And yet it's still being seen as a major confrontation between the mathematical community and the AI industry because Buckmaster is a prominent member of the community and has taken pains to follow mathematical norms while OpenAI has not (with the most flagrant violation being the insistence on removing Alpöge from authorship, simply because of corporate affiliation), and is more nakedly threatening human ownership with capital (the OAI blog post says the final result involved 10,000 concurrent agents).

Regarding credit assignment when LLMs are involved, the mathematical community has organized around some rough principles. Gowers has some thoughts on his blog (https://gowers.wordpress.com/2026/07/26/thoughts-about-the-l...) about how explaining a result may be more deserving of credit than producing it. It looks like Buckmaster and Alpöge were taking their time in understanding their results and writing them up when OpenAI forced them to publish their work-in-progress. At the same time OpenAI has published their own writeup but it's not really clear to me how involved humans were.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#626

> This is a a Deep Blue-Kasparov moment. I guess this is true in more ways than one. Kasparov famously accused IBM of cheating during the match, by spying on his preparation (edit: though the main cheating accusation was live human intervention during the games, on top of IBM downplaying the heavy human involvement behind the AI, which also mirrors this situation)

Deep Blue beat Kasparov fair and square. Kasparov was a bit of a bad sport at the end of the match, though the reasons are understandable. He was at the top of the human chess world. He wasn't used to losing, and he took it badly.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#627
post #622

Earlier quoted context omitted.

Perhaps I am underestimating the size of these frontier models then, how big of a cluster do you think would be required to serve a university?

I don't think you're taking this thread seriously enough to respond in detail. A university can be tiny, or it can have thousands of researchers. Some of them might want to work on small models, and others on huge frontier models. Multiple groups training their own models at the same time. Combined with the storage, networking, power, and redundancy, you basically need large data center scale. And at that point you'r…

I'm aware cluster sizes vary, I was mostly wondering how long much compute you need to serve a frontier model (since you seem knowledgeable on this topic). For instance back of the enveloppe Kimi K3 fits in ~24 H100, so my naive first impression is thats its not out of reach for a university-sized cluster to serve a few instances. I agree it may not make much economic sense I'm asking out of curiosity.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#628

Earlier quoted context omitted.

OAI doesn't need to mention Buckmaster's name directly in a prompt. They just need to select a basket of sessions that is guaranteed to contain Buckmaster's and then direct the LLM to attack only a specific method/angle. This is trivial to do while maintaining plausible deniability about not using his work.

what reason do we have to believe that they did this? both things were proved by AI, isn't it logical that they could have very similar approaches? it is common that multiple people essentially simultaneously prove/invent the same thing I see zero evidence of wrongdoing

> what reason do we have to believe that they did this?

The culture at OpenAI being systematically revealed by Apple’s lawsuit, for one.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#629
post #83

Drama/accusation summary: - Aug 15th: Tristan Buckmaster & Levent Alpöge make progress on a few important math problems, "finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler." - they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help le…

It's specifically the last two bullet poitns - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. - OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) — but…

> These two bullet points are extremely suspicious if you were honest

I don’t think we can consider these accusations separately from the evidence being unveiled about OpenAI’s culture by Apple’s lawsuit. These guys seem to openly embrace the strongest interpretations of “good artists copy, great artists steal.”

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#630

Earlier quoted context omitted.

I think you may be underestimating how difficult a text search over their data is. They may have to build new mechanisms to do this. And what you really want is also an attribution of how much of a contribution a given corpus made which is a much harder question to answer; a single appearance of a chat probably has very little impact on the inference performance at this time unless it’s been explicitly preferenced so…

If they literally can’t audit training data for a given model, they shouldn’t be operating.

> If they literally can’t audit training data for a given model, they shouldn’t be operating

They can operate. They shouldn’t be claiming credit for discovering anything.

Post reply on HN