Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

811–820 of 871 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#811

Earlier quoted context omitted.

that is not related to the accusation that they literally stole Buckmaster's work, which seems baseless I don't agree with the oil claim analogy. this is knowledge, freely given to the world. not something hoarded by a corporation

but a very large part of the whole model was trained on work in a manner the authors didn't consent to, the "for research purposes only" datasets of the entire Internet, etc and you can argue this is "fair use" or whatever, not the point now, the point is that it definitely makes those accusations no longer "baseless". in addition, it is not given freely to the world, it is the knowledge of the Internet/WWW being sol…

Humans do this all the time. You watch a YouTube video and subconsciously choose the same colour palette. They hear a rumour that it's a solved problem, i.e. they were pointed in the right direction that is all.

I mean some companies glean insight into new products merely by asking other people what they do for a living.

People need to get over this ownership thing, it's being taken too far. Humans benefit from the efforts of others simple as that.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#812
post #643
post #83

Drama/accusation summary: - Aug 15th: Tristan Buckmaster & Levent Alpöge make progress on a few important math problems, "finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler." - they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help le…

I don't understand the desire to remove Levent. "Off the clock, when Anthropic engineers want to break new ground, they use ChatGPT" sounds like a great ad.

All these small details will be forgotten in less than a year. But association "Navier-Stokes -- OpenAI" will be part of the history. In my opinion, they are thinking about long term PR here.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#813

People here do not seem to be considering the second-order effects of these series of events. No academic institution or enterprise will trust OpenAI, Anthropic or any other non-local AI model with their core IP. There will be severe restrictions on what employees at these companies/institutions can share with AI services even from their personal accounts. (Or I am just overthinking it)

Many academics and grad students I know have closed source their in progress work, and started being really careful about what they chat with LLMs (or using local ones) because of the drama around this. No one wants four years of their life getting sniped by ten million dollars worth of tokens.

Or, if you're using, you need to push a pre-print relatively soon after AI help.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#814
post #724

Earlier quoted context omitted.

This is being reported as OpenAI wanting to strip an Anthropic employee of academic credit for the work they did. What the OpenAI person involved is claiming is that they wanted the outside researcher(s) to put their name on OpenAI's work: to headline OpenAI's publication of what they earnestly believed to be an independent result. If true, that's generous and beyond the level of generosity one should expect. Extendi…

> What the OpenAI person involved is claiming is that they wanted the outside researcher(s) to put their name on OpenAI's work > If true, that's generous and beyond the level of generosity one should expect "We highly likely stole your work, and threatened you with 'this is bad for your career' and we refuse to acknowledge any work by your collaborator just because he works at a competitor, but we are so so so so gen…

The two proofs are structurally very different, and don’t even prove the same conjecture. It’s becoming very clear that OpenAI did not steal anything here.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#815

Earlier quoted context omitted.

I work in catastrophe risk modeling and it's a multi billion dollar industry. We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do. There's a lot of pressure on AI adoption so the company has partnered with various tech companies to build intelligent systems on top of proprietary data and m…

> We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do. I feel this is exactly what will happen as they cause all sites to go closed source to protect their intellectual property and the AI companies offer only biased information. They are replace the business on internet model by bankrupti…

This is the business case already. And has been the case with tech companies for a long time. Your phones built in photo manager replaced a lot of what Photoshop does.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#816
post #57

So, leaving aside the idea that OA might've used data from the researchers Codex sessions: Do I understand correctly that the internal OpenAI work on the problems was probably started after they heard Alpoge and Buckmaster had made process by using their models? And they used the publicly available info about the researchers past work to prompt their models? If compute is cheap, and the difficult thing with scientifi…

It's honestly unsurprising and not a problem that they do this in my view. The problem really starts when you start taking credit for work that they would've achieved. Like if i go to a talk on unfinished work, it's not really unethical for me to think about the problem--it's a problem if i scoop the authors but these problems can often be solved by collaboration or proper crediting and timing--IN MY VIEW

The difference is how credit and attribution works. And whether we feel it’s being laundered through models.

And also whether the AI moon laser pointed at your problem is just going to be the thing that writes the final conclusion on ten years of your work.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#817

Earlier quoted context omitted.

Frankly, I don't buy this difficulty argument. They know which model was used to come up with that particular idea. A text search over the corpus of user data used in the training set can only take so long.

I worked in the tracing and tracking all the thousands of data sets that got tweaked and permuted and changed hands between thousands of researchers and data engineers at a major lab. The data that goes into training runs is permuted so much from the OG data that tracing the lineage is not trivial (dramatic understatement). And the difficulty is harder than just the extreme scale of text searching. but also explodes…

This sounds like data laundering in effect.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#818

Earlier quoted context omitted.

It's specifically the last two bullet poitns - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. - OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) — but…

If that part of the PDF is true that’s so disgusting, psychopathic behaviour > I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.” Some time later Lev…

This kind of coercive, threatening rhetoric is really not surprising at all. This is how tech companies operate.

What stands out to me is the naive lack of operational security on the part of academics, who should know better than to touch this SaaS crap with a ten foot pole.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#819

Earlier quoted context omitted.

Just replace the model with a human student. "Training" on textbooks => fine "Training" with unpublished notes from another professor, then publishing something on that exact topic with a similar approach without giving any credit => extremely questionable.

Presumably the professor voluntarily provided the notes in this analogy. I think the student would also be expected to cite the textbook if building off of it directly. In contrast, humans are generally not expected to cite "general inspiration" or what have you. So if we're to apply human standards, and assuming that the model was trained on the relevant work, it would only be plagiarism if the model directly built…

voluntarily is stretching it for an opt-out

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#820
post #364

OpenAI's release explicitly says No. But then also caveats that with "we cannot rule out that de-identified data derived from their usage of our products" impacted things. What's most striking to me, and what may or may not be true, is the "we cannot rule out" bit. "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user dat…

That was also striking to me too, the 'we cannot rule out'.

I also do not believe it. They can, trivially, rule out their model being trained on if

More to the point -- of _course_ they know what data their model is trained on and the lineage of that data, even if it ended up being anonymized and they cannot identify which precise user.

Post reply on HN