Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

511–520 of 844 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#511

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them,…

Nooo don't be mean to the mass plagiarism machine that since day 0 has stolen anything and everything they can get their hands on!!!!! We need to give unscrupulous corporations who time and time scam, cheat and lie their way through everything infinite grace to fuck up the planet!!!! Think of the IPOs!!!!!!

Re: More questions about whether researchers can trust OpenAI with unpublished math

#512
post #137

Earlier quoted context omitted.

That’s one perspective. I just view it as a thing that can brute force and produce outputs - that it has no way of ‘knowing’ - but doesn’t need to since it’s just running off of probability. No human can compete in that contest. But no llm can compete in the contest of ‘understanding’ and application in the real world - which is where 99% of the value is. I’m very pro AI long term btw but I’m not blinded.

AI needs humans to encode ideas in words. It needs those ideas to span the space of possibilities of, say, Navier Stokes. Then AI can be, as you say, a terrifyingly effective way to search that space. But when the building-block ideas are still being formed, I'm not sure that AI is good at forming them.

COrrect and this is how labour displacement happens.

There are many actions being performed today that can be nicely packaged.

Im already working on such a project.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#513

Earlier quoted context omitted.

First, OpenAI is not claiming that the model wasn't trained on those sessions. What they've said is “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.” and “We did not use their prompts or proofs to prompt our models or direct our agents.” and “While unlikely, we cannot…

> It's not obvious to me that's an unethical thing to do In terms of work in mathematics, something I personally would not do based on ethical grounds would be to hear a rumor that some researchers are taking a certain approach and may be nearing a solution, use a model that was possibly contaminated with intimate knowledge about that approach (though later they investigated and think it wasn't), and then commit mill…

I heard they also tried to strong-arm them into removing the name of their collaborator who happened to work at a different company (Anthropic)...

I haven't looked into it myself, but if true, that seems incredibly scummy.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#515

Earlier quoted context omitted.

"Discovery" is not a goal in itself. I could launch a project to find out how many people in the United States have names such that if you assign numbers to every character and then sum the values, the sum works out to 72. It's discovery, but it's useless unless it has some higher goal. The labs are attacking these problems as a demonstration of capabilities, spending more money on the demos than any mathematician wi…

Right, mathematicians care about clout and tenure, which is a much higher purpose.

Yes, blame them for seeking out an upper middle class lifestyle with a relatively standard home in commuting distance of their place of work and dedicating the rest of their life to teaching mathematics to new generations of people. How vain a pursuit.

After all, the ascetics at openAI are having to make do with half a million total comp.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#516
This smells of extreme 'cope'. Am I really supposed to believe that all of these problems could have been solved, were right about to be solved, etc. But it just happens they are all getting solved now when AI is getting really good at Math...

Re: More questions about whether researchers can trust OpenAI with unpublished math

#517
post #137

Earlier quoted context omitted.

That’s one perspective. I just view it as a thing that can brute force and produce outputs - that it has no way of ‘knowing’ - but doesn’t need to since it’s just running off of probability. No human can compete in that contest. But no llm can compete in the contest of ‘understanding’ and application in the real world - which is where 99% of the value is. I’m very pro AI long term btw but I’m not blinded.

But it’s not brute force if it’s looking over everyone’s shoulder Brute force would have been solving Navier-Stokes in 88 hours after plagiarizing all known 20th century math When it needs to snoop live on what the actual mathematicians are working on that’s something else

No its happening whilst the human is working with it. The new inputs provided become part of the brute-force. This is what Scam Altman means by 'self-recursive'.

Trust me I've seen it happen to myself. I no longer trust ChatGPT.

I can see right through his act. Altman is one devious f8k.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#518

Earlier quoted context omitted.

> I've been wondering whether AI really is improving rapidly at open problems or we're being fooled. I think your suspicions are warranted and your explanation seems plausible. If better training data is the reason here, it would still be a case of the models doing something that is in and of itself super useful! The models really can take that data and distill it into solutions for similar problems faster than human…

>> If better training data is the reason here, it would still be a case of the models doing something that is in and of itself super useful! The models really can take that data and distill it into solutions for similar problems faster than humans can. This is great! It's perhaps great in the short term although it's not very clear who it's great for. I'm not sure mathematicians find it all so great, I mean. In the l…

I was pretty depressed when I read about what happened with Navier Stokes this morning. The Clay Math prizes were a significant motivation through my math career, and I know a lot of computer scientists and physicists that feel similarly. I didn't think I was gonna resolve P vs NP or the BSD conjecture, but I did really research that felt like I was working towards something incredible. What is the younger generation left with? Hey kids betcha can't resolve the Collatz conjecture, our superintelligence can't either! Still pretty depressed about it, to be honest. Intellectualism is dead. We can return to happy agrarianism, I guess. At least the AI doesn't wanna eat my snap peas.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#519

Earlier quoted context omitted.

When you say "confidential in-house version", what are you referring to? Local models? Bedrock deployment with "guardrails"? A different thing?

Enterprise Agreements can have binding terms for this. When I launch the ChatGPT desktop app, and open the options pane it says "Corpname data is not used for OpenAI training". I would expect academic institutions to require equivalent contractual terms.

Sure but they could also rewrite your data to create synthetic reconstructions and many academics, sign up for their own accounts.

For example, at school they can have an agreement with Gemini, but the student / academic could have bought an individual pro subscription to any other model provider.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#520

Earlier quoted context omitted.

Good question, this was very well known. Do you have an answer?

There is no fine-print (let alone a loud banner) on the chat thread page that tells me my prompts can be used for training.

Every single internet connect piece of software there is probably collects telemetry at this point. Why would this be any different? You know google logs your search data as well right? Not just the companies scan it but law enforcement too.
Post reply on HN