Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

721–730 of 849 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#721
post #708
post #632

Earlier quoted context omitted.

But by OpenAI's telling they heard a rumor that the problem had already been solved. So they reached out to the other researchers as an attempt to share the credit, and in fact have at least one of them be the lead author (which is when they found out the AI had solved a broader problem than the researchers.) Seems pretty ethically palatable. I suspect the main reason the community is not receiving it well is largely…

as a developer that had a brief career in academia, i don't think your last comment is right at all. 99.9% of what i work on as a webdev, even if it's challenging and unique at the margins, is not really novel. concerns about job security aside, i don't really think of an agent as stealing my ideas because it's good at writing CRUD APIs. collaborating with ChatGPT on a novel solution to an unsolved problem, getting 9…

> concerns about job security aside

But that is exactly what I'm implying is the core reason, whether people realize it or not.

I totally agree that the vast majority of software dev is not novel. I have even made several comments to that effect. The same can be said for a lot of creative work as well. Yet many, many devs and creators are very unhappy with AI, and a lot of their complaints are variations on accusations of plagiarism.

And note, I am not saying it is wrong, it is completely understandable, but we need to be clear about where this turmoil is coming from.

If I were in the same situation as these researchers, I would publish all pertinent research work and chats so that the rest of the world can see how close the model's work is to my own. It's been scooped anyway, so there is no reason to keep it private.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#722
post #706

All of these accusations could be true. But there's also no way for a company to casually claim "No, we did not train on your data", without verifying all the knobs the user might have turned to enable or disable data sharing. I just don't understand getting the pitchforks out because a company did not give an answer immediately. And the effect such data entering training would have affected the output is even less c…

The pitchforks are out because even without that part, it's still a scumbag move to try to frontrun the mathematicians who were working on this for years after OpenAI heard that they were close to releasing their results. Just identifying that one of these problems is solvable takes a lot of work. The only reason OpenAI got this result is because the mathematician shared with colleagues that he had made significant progress and was close to solving it, and OpenAI could not accept that so they decided to throw tens of millions to make sure it doesn't happen without them getting all the glory. Notice that their paper doesn't even have an author since they're probably all aware of how awful that would look, and no one wanted to take on the shame. They probably also knew that the paper was trash and no one involved could understand it, and didn't even cite many of the people who contributed to all of that knowledge. It's just a disgusting act any way you slice it, even without training on the prompts or the nasty communication by the OpenAI leaders.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#723

[dead]

This. These platforms are asking to be trusted with unprecedented amounts of the public's data and, unprecedentedly itself, the public's reasoning and decision-making. It's an awesome responsibility that requires a singular approach that smaller platforms with less responsibility don't necessarily have to devote resources to. OpenAI, Anthropic, Google, Facebook, they're the big dogs. They can't do the things the smal…

Nice analogy! This seems to me the main difference between Google and the other companies you list -- In regards to LLMs Google is the only one mostly acting like an mature, established organization and not rushing to market as soon as possible. For their responsibility they often get labelled as having fumbled some imagined race. Maybe they actually do just lack the talent to make better models but it's not like they aren't making advances in other areas of AI.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#724

Earlier quoted context omitted.

I was referring to the overall pattern of apparently sniffing around for recent mathematical progress then setting the AI on it to see if the problem is now easy enough to solve (if you have the money). Terrance Tao has lamented this practice as being unhelpful for mathematics, and likely to lead to humans working in private to avoid this. Tao has also noted that many of these AI math proofs don't really help mathema…

> has lamented this practice as being unhelpful for mathematics A related point is that the actual solution approach is never revealed. What was the role of humans guiding the agents ? was it fully autonomous ? etc. It is in the incentive of the AI labs to trump the powers of the LLM, but in practice it is humans guiding the agents on the overall approach, This is never admitted. For example, in the announcement on N…

I don't think it's as bad as that sounds; in math people work all the time with conjectures they aren't sure if true, and work out a lot of other interesting math based on whether it is or not. Something like Turing's Oracle machine gives lots of interesting math just assuming one could exist, even if it couldn't. It may be that there are things proved we can never come to a human understanding of, but still keep getting interesting math that relies on it that has aspects we can appreciate and enrich our knowledge from.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#725

Earlier quoted context omitted.

That's also incorrect. My understanding is that they asked the independent researcher to improve OpenAI's AI generated proof and be the lead author of the paper to publish OpenAI's result. This is the paper where they did not want the Anthropic employee collaborating. Not their work.

I think both Seb and Sam have said that it would’ve been simpler if the coauthor hadn’t worked at Anthropic so they’ve largely admitted they didn’t invite the collaborator as a coauthor because it would’ve look bad to have an Anthropic employee on the paper.

seems pretty short-sighted - "our models are so good that even our competitors use them for the most advanced tasks" is pretty powerful marketing

Re: More questions about whether researchers can trust OpenAI with unpublished math

#726
post #360

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

On your second point: there is a more plausible explanation which David Bessis calls the "overhang". The short version is that there is a large amount of relatively low hanging fruits in mathematics, because no human has broad enough knowledge and enough time to try them all. AI is not constraint by that, and therefore can systematically pluck all those low hanging fruits. Quote: "The Overhang consists of the unreali…

The overhang, being defined as the Cartesian product of existing knowledge — randomly combining existing knowledge.

(I mean actually randomly, not asking an LLM to do the randomness.)

Most of the output would be incoherent (like many dreams), but occasionally you would get a gem.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#727

Earlier quoted context omitted.

Both can be true: 1. OpenAI couldn't have solved the problem without the researchers' private data for training. 2. OpenAI models can solve math problems

Anthropic isnt getting enough scrutiny for their unprofessionalism: 1. Anthropic employee working on monumental problem but didnt receive/ask for the full backing of the company's resources 2. May or may not be mixing unreleased Claude output with Codex without zero data retention agreement 3. Victory lap on Twitter and giggling around the city before they finished the job, sparking rumors for competitors

How dare employees do something without asking for the full backing of the company's resources. Incredibly unethical!

Re: More questions about whether researchers can trust OpenAI with unpublished math

#728
post #649

Earlier quoted context omitted.

[dead]

Was it? They've been dishing out cheap access specifically to researchers give over lmao. The researcher's got lured in - they need to accept they got played TBH. Altman is certainly more devious than Amodei - he's shown that time and time again. PG was right about he said about him. Every entity on earth should see it as a kill shot: be careful what you put in the models. None of your information is safe.

i'm not sure why there's a need for comparison here. both are not saints and you shouldn't trust any of them anyway

Re: More questions about whether researchers can trust OpenAI with unpublished math

#730
post #706

All of these accusations could be true. But there's also no way for a company to casually claim "No, we did not train on your data", without verifying all the knobs the user might have turned to enable or disable data sharing. I just don't understand getting the pitchforks out because a company did not give an answer immediately. And the effect such data entering training would have affected the output is even less c…

The pitchforks are out because even without that part, it's still a scumbag move to try to frontrun the mathematicians who were working on this for years after OpenAI heard that they were close to releasing their results. Just identifying that one of these problems is solvable takes a lot of work. The only reason OpenAI got this result is because the mathematician shared with colleagues that he had made significant p…

> They probably also knew that the paper was trash

Doesn't matter. This forum used to celebrate "because you can" with no riders. And solving a Millennium Prize problem is among the biggest stages for Because We Can.

Now we're saying there are some qualifiers attached to it, such as (1) only if not done by companies with a lot of money, (2) only if it is inconsequential.

I agree with some of what you're saying, but like everything else it isn't black and white. Maybe some day, someone will improve some particular treatment because we can.

Post reply on HN