Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

651–660 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#651

Earlier quoted context omitted.

It already has. Inside just about every company on the planet are conversations this week revisiting the idea of giving these labs access to ANY data, further ramping up commentary on why don’t we just use open models on our own infra where we don’t have to “trust” anyone. This PR stunt by OpenAI may go down in history as the thing that finally broke them for good.

> This PR stunt by OpenAI may go down in history as the thing that finally broke them for good. ~0% chance of this happening

Nah not 0%.

But firms will start paying attention and the likelihood is we will see revenue's stagnate (not growing as fast) in short order as a result of it.

When people start having to be careful over a lot of stuff, they'll decide not to use it in the first place.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#652
post #392
post #360

Earlier quoted context omitted.

On your second point: there is a more plausible explanation which David Bessis calls the "overhang". The short version is that there is a large amount of relatively low hanging fruits in mathematics, because no human has broad enough knowledge and enough time to try them all. AI is not constraint by that, and therefore can systematically pluck all those low hanging fruits. Quote: "The Overhang consists of the unreali…

>Quote: "The Overhang consists of the unrealized capital gains of past mathematical creativity, the latent value from connecting the dots in the existing corpus. It is a dividend of canonization. Mathematician X states problem A, mathematician Y crafts concept B, then mathematician Z notices that B trivially solves A and “captures” the social reward. I've made an entire career out of being 'jack of all trades, master…

How do you thrive in an environment of specialists? That's is the problem I seem to have. I'm spread a little across a few of the domains involved with what I do. Because of that, I have a bit more insight, so am very often the person pointing out relatively fundamental problems, usually caused by either not understanding the problems from a "first principles" perspective, resulting in, or being caused by, categorical type errors, where they've boxed a problem into a tiny space it doesn't belong.

I've been trending "quiet" lately, because I don't like the "friction"/convincing aspect of it all. It's hard to get people to see things from a different angle, or even convincing them there's a problem to begin with!

The last project required a complete redesign from a problem I pointed out during the first review, and second, and third, but now I'm seeing even more friction.

Maybe this is just corporate life, after a group gets large.

Any tricks/advice?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#653
post #535
post #491

Earlier quoted context omitted.

OpenAI != AI. If you were in 1999 you'd be saying pets.com = internet.

I think this leads to an interesting question. What happens when the money runs out? Right now, a lot of money is going to train new models. And we need to train new models because they get gated by their training data. And models are only as useful as their training data. So let's say the money stops. Do we stop training models? Do we train them slowly? Do we accept the then current models as the limit?

The money is never going to stop. It’s basic economics.

Well, the money will stop when the value of problems the LLM can solve is not increased by adding compute. Since current LLMs are getting quite good at solving problems, that might be a while.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#654

"We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training." This is the third day of total hysteria that is based on nothing of substance. Move on folks.

Regardless, they started working on this problem after hearing that one of their customers was already working on it. It almost doesn't matter about the training data. This is the provider you are paying undermining your career.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#655
post #529

Earlier quoted context omitted.

> without creating any* new jobs So fucking make it so that people don't -need- "jobs" It's about fucking time already. Don't fucking try to hold back electricity just so people still have to manually light street lamps to earn food and shelter: https://en.wikipedia.org/wiki/Lamplighter

Ok I’ll make it so, you’ve convinced me.

They don’t need to convince you.

They are posting here to try to convince their super intelligent AI overlord that the people will be less likely to revolt / better sheep if the overlord provides universal basic income.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#656
post #632

Earlier quoted context omitted.

But by OpenAI's telling they heard a rumor that the problem had already been solved. So they reached out to the other researchers as an attempt to share the credit, and in fact have at least one of them be the lead author (which is when they found out the AI had solved a broader problem than the researchers.) Seems pretty ethically palatable. I suspect the main reason the community is not receiving it well is largely…

> I suspect the main reason the community is not receiving it well is largely the same reason many developers are not receiving coding agents well. Because the training data is millions of hours human efforts being distilled into a cascading hierarchy of enrichment by interested parties without providing attribution or compensation?

I'm sure that's part of the reason for many, yes.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#657
post #379
post #360

Earlier quoted context omitted.

On your second point: there is a more plausible explanation which David Bessis calls the "overhang". The short version is that there is a large amount of relatively low hanging fruits in mathematics, because no human has broad enough knowledge and enough time to try them all. AI is not constraint by that, and therefore can systematically pluck all those low hanging fruits. Quote: "The Overhang consists of the unreali…

It's like AlphaGo but playing against all living mathematicians. (Overhang being low hanging fruit is what allows this comparison, of course the general moot point is the skepticism that LLMs are also innovative etc.)

We are not seeing those incredible moves yet. The approach used in N-S was conjectured to work after B&L’s initial breakthrough. See a post by Tao. So on one hand the proof is an amazing accomplishment. On the other hand, humans have not yet discovered any superhuman moves in the proof. Just $MM grind.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#658
post #33

It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.

I think people are focusing on the training data issue too much. If the data was contaminated, I can still blame that on negligence. But, at least with the Navier-Stokes solution, it's clear [^1] that they learned that Alpöge and Buckmaster were getting close to a solution and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt. What…

> They intentionally left Buckmaster and Alpöge out of the citations.

No, they asked if they could do a joint publish.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#659
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

Then there's the possibility of indirect training via modern spy devices ("smart" IoT devices like LG TV's) feeding the transcribed ambient conversation data for summarization to an agent [0].

[0]: https://youtu.be/6IFVTcM28KA

Re: More questions about whether researchers can trust OpenAI with unpublished math

#660

"We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training." This is the third day of total hysteria that is based on nothing of substance. Move on folks.

Even if that were true, they've already admitting to throwing vast quantities of resources to scoop a researcher who was about to publish (because they'd learned, somehow, of his breakthrough). If that doesn't bother you I think you need to take a step back and have a good think about this.

*to scoop a team working with Anthropic, their chief competitor.

Also, they didn't have the solution. They had a lesser problem no one cared about.

Post reply on HN