Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

201–210 of 848 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#202

So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training? How do people become that trusting? The phrasing itself is guilt tripping

Personally, I wouldn’t assume it was lying. To me, dark patterns (like manipulative wording) imply that: 1) someone in a governing body, or someone in the organization, e.g a designer, ethicist, lawyer, developer, etc. has successfully argued that users should be able to avoid something that they determine is not in their best interest. And also: 2) someone in the c-suite or marketing has decided to mitigate that thr…

> To me, dark patterns (like manipulative wording) imply that:

Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues with lies and ends with lies.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#204

I've been wondering whether AI really is improving rapidly at open problems or we're being fooled. - OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay - Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2] - But researchers will typically work on open problems. A researcher who is using Co…

A mental model I was thinking about was - I remember when Travis Kalanick was talking about using the chatbot to discuss “vibe physics-ing” on the all-in podcast.

And like - I think there’s a presumption you could make that AI models could overfit to asymptote towards just the capabilities and knowledge we currently have.

And that would be amazing! And crazy useful. And there are probably a whole world of complex problems that remain unsolved because they’re adjacent to knowledge we have but they haven’t been invested in.

But can a human reliably tell the difference between “can do 99.999% of the things we currently know how to do which includes a small subset of things we didn’t know we had the capacity to do” and “super intelligent math and science research pushing the frontier of what we know”

A physicist that knows all the things we currently know in excruciating detail feels like it should be able to make the leap beyond the frontier.

But since these are computer models it might just be that it can ride that line extraordinarily well while the line remains firm.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#205

Earlier quoted context omitted.

"Discovery" is not a goal in itself. I could launch a project to find out how many people in the United States have names such that if you assign numbers to every character and then sum the values, the sum works out to 72. It's discovery, but it's useless unless it has some higher goal. The labs are attacking these problems as a demonstration of capabilities, spending more money on the demos than any mathematician wi…

Right, mathematicians care about clout and tenure, which is a much higher purpose.

Are you saying that mathematicians are the bad actors here? Compared to Sam Altman spending ungodly amounts of money to upstage them ahead of IPO?

I care about paying my bills and job security and peer recognition. That's a normal human thing to do, not some vice. You don't?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#206
post #137

Earlier quoted context omitted.

> they could be significantly piggybacking on human progress, This is AI in a nutshell, its a plagiarism machine. An abstraction layer between vast amounts of stolen human-generated data that filters out the liabilities and accountability for that original theft. Its an IP laundering system.

That’s one perspective. I just view it as a thing that can brute force and produce outputs - that it has no way of ‘knowing’ - but doesn’t need to since it’s just running off of probability. No human can compete in that contest. But no llm can compete in the contest of ‘understanding’ and application in the real world - which is where 99% of the value is. I’m very pro AI long term btw but I’m not blinded.

But it’s not brute force if it’s looking over everyone’s shoulder

Brute force would have been solving Navier-Stokes in 88 hours after plagiarizing all known 20th century math

When it needs to snoop live on what the actual mathematicians are working on that’s something else

Re: More questions about whether researchers can trust OpenAI with unpublished math

#207
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

There's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem". I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).

But who gets credit then? Every mathematician who's work was read by an LLM during training? By that logic, we should put every published mathematician's name on the authorship of this paper. Sure, this guy should be higher up the list, but everyone's name should be on it by standard academic convention.

But this gets back to the original "who owns the LLM output" and "can you train models on the internet" argument that's been raging for years.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#208

I think this is stupid, for three reasons: 1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods. 2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether the…

> in this case too they didn't actually have the solution Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data. I would almost expect training to overweight conversations with novel scientific and mathematical implications. > the only reason they can't definitively say no is that for privacy reasons They could 100% d…

You quoted the OP saying "in this case too they didn't actually have the solution" and responded with the totally unrelated, "Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data."

Neither of the researchers insinuating that their ideas were trained on had the actual solutions. This means the model could not have "stolen" the final solution from their data. At most, it could have built upon their work in the same it builds upon any other training data, though that is also questionable speculation.

>They could 100% definitely say no, if they know they did not train on user data.

No one anywhere has claimed that "OpenAI does not train on user data." OpenAI has always said that it trains on user data.

>They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.

They started racing towards a solution after they heard (incorrectly) that Anthropic had a solution; I agree this is poor sport but the "after one researcher enquired about whether they are training on their conversations" claim is false. The enquiry happened after OpenAI had obtained the solution.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#210
post #137

Earlier quoted context omitted.

> they could be significantly piggybacking on human progress, This is AI in a nutshell, its a plagiarism machine. An abstraction layer between vast amounts of stolen human-generated data that filters out the liabilities and accountability for that original theft. Its an IP laundering system.

That’s one perspective. I just view it as a thing that can brute force and produce outputs - that it has no way of ‘knowing’ - but doesn’t need to since it’s just running off of probability. No human can compete in that contest. But no llm can compete in the contest of ‘understanding’ and application in the real world - which is where 99% of the value is. I’m very pro AI long term btw but I’m not blinded.

AI needs humans to encode ideas in words. It needs those ideas to span the space of possibilities of, say, Navier Stokes. Then AI can be, as you say, a terrifyingly effective way to search that space.

But when the building-block ideas are still being formed, I'm not sure that AI is good at forming them.

Post reply on HN