Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

841–850 of 873 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#841
It's rather convenient for OpenAI that user logs are de-identified before being fed into training, so they can say "there's no way for us to check if we plagiarized our user's work, because user data is private".

It's also difficult to imagine any competent AI researcher would overlook the possibility of training data leaking into the test, especially given that they know the users have been using their model to work on the same problem, and that they jumped on the problem after hearing rumors of the breakthrough.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#842
post #581
post #536

Earlier quoted context omitted.

> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case. The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor: 1) the agents spin for days and produce too much output to review 2) using LLMs to process that output skips many important details Ergo, the agent could likely decide it would l…

> The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor It requires humans to verify what agents have done. Weird

which is rapidly becoming a game of steganography cat and mouse

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#843
post #724

Earlier quoted context omitted.

> What the OpenAI person involved is claiming is that they wanted the outside researcher(s) to put their name on OpenAI's work > If true, that's generous and beyond the level of generosity one should expect "We highly likely stole your work, and threatened you with 'this is bad for your career' and we refuse to acknowledge any work by your collaborator just because he works at a competitor, but we are so so so so gen…

The two proofs are structurally very different, and don’t even prove the same conjecture. It’s becoming very clear that OpenAI did not steal anything here.

Bullshit.

Ain’t no way you read one 100 page paper let alone two with enough understanding to make such a claim.

You ought to be ashamed of yourself.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#844
post #312

Earlier quoted context omitted.

I'm stunned that people are taking this accusation as a fact. OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business. There are things that Buckmaster alleged and things that he specu…

> The entire training data thing is speculation. I think it's safe to assume AI labs DO train on your data and it's very hard to prevent that. I've just checked my inaptly named "Help improve our AI models" toggles. The toggle on the Claude settings had magically turned on. I asked about how this can happen. Claude says they show re-consent modals when terms change, and it is a "real and fairly common pattern" to re-…

You can opt-out of training in Gemini on personal plans, but it disables your chat history, just to be vindictive; there is no technical reason and the other companies don't do this.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#845
If you'd like to explore what the Navier–Stokes blow-up construction looks like visually, I vibe-coded an interactive 3D visualization based on the published result for fun :) Demo: https://minfx.ai/navier-stokes/ Source code: https://github.com/minfx-ai/navier-stokes-blowup

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#846
My take is that, if you really care about humanity and how LLMs can be utilized to benefit mankind, you should try to solve hard problems, try to endeavor, instead of burning money just to publish a paper ahead of your human competitors, especially when they already claimed that the problem had been solved and were just spending time formalizing the publication. Maybe legit but still disgraceful. OpenAI should at least disclose how they utilized their internal models to solve the problem, to show human how we can empower ourselves with AI. Now they just seem like colonizers who are desperate to plant a flag on every land within eyesight.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#849

Earlier quoted context omitted.

Agree, I think the practice is also very clear from the overall strategy of AI-companies and their ToS: Scale with subsidized pricing as fast as possible to gain more user-data for training --> Own the better model --> scale pricing. Scanning social media (e.g. Twitter, Reddit) posts only give a glimpse into the thought-process, chat logs on-scale give you the actual process in machine-readable format. There's a reas…

> - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. The question, for AI customers, is when they build products using services of AI-companies, would AI-companies engage in theft of customer data for use in training?

If you still had that question, you can answer it now.

But honestly... "Will the company that was entirely built over illegally acquiring data use some data that is legal to use and is right on their front, or will they not do everything they reserve the right to do?" is a really bad question for one to even ask.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#850

One of the best takedowns of samas/OAIs spin on this I've seen is this post on X: https://xcancel.com/BetterCallMedhi/status/20974723362578637...

the first "sentence" just seems to continue indefinitely into word salad

[flagged]
Post reply on HN