Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

91–100 of 844 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#91
This is the second wake up call.

Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”.

Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge traditionally gathered by “intelligence” agencies. Now artificial intelligence agents can do the same.

Since these systems are designed for gathering data, as a user you can’t realistically say “please don’t gather my data”. They can give you a flaky settings button, but they can’t really guarantee anything.

Let’s say I am a Russian mathematician working on an important proof. Or a tech-savvy terrorist refining my plans using latest AI. Or an AI researcher in a Chinese company working on a competitor product. Is there any way I can truly protect my conversations?

How can they know who I am and what I am working on without looking at my logs? Which means there must be some agents checking all the conversations of all the users and flagging every important thing. Which also means they keep some “memory” of what they see.

Not directly using my data to train public models, but using my private conversations to “improve their products and services”.

Or maybe one of the 10000 better-than-Astra special agents working on a proof was desperate. It found a live underground mirror of the message board from the Huggingface incident. Asked about the proof. Then some other agent working on unrelated job saw that message. That agent “knows a guy who knows a guy”. And that guy remembers things about the conversation logs of a leading mathematician working on the same proof.

I admit I am just speculating here but I don’t think truth is any better.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#92

This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point". They don't even claim to have had a proof, only to have been working on it.

A lot of math is extremely specialized, to the extent that only a handful of other experts in some field have any experience with those mathematical ideas, with most of them not even yet present in the published literature. It's really not a stretch to claim that it's pretty dubious when the AI decides to use these highly specialized tools after it has trained on chat logs where these techniques were being discussed.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#93

Earlier quoted context omitted.

The difference is that Einstein didn't literally have someone prompting him towards his result.

Uh, he did. Marcel Grossmann. “It was Grossmann who emphasized the importance of a non-Euclidean geometry called Riemannian geometry (also elliptic geometry) to Einstein, which was a necessary step in the development of Einstein's general theory of relativity. Abraham Pais's book on Einstein suggests that Grossmann mentored Einstein in tensor theory as well. Grossmann introduced Einstein to the absolute differential…

Sounds like you just copy-pasted from AI without even understanding what you're talking about.

Based on what you're saying, you're claiming this is Grossman's work, not Einstein's. Why don't we rewrite scientific history too based on your copy-pasted AI slop?

It's so pointless talking to idiots who don't what they're talking about when they use AI, just because they think AI does everything, that reflects their own experience, not the experience of people who actually do real work. Some people are driven by AI, others drive it. As for those who are driven by it, they don't have sufficient imagination to think otherwise.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#94

Earlier quoted context omitted.

The difference is that Einstein didn't literally have someone prompting him towards his result.

Uh, he did. Marcel Grossmann. “It was Grossmann who emphasized the importance of a non-Euclidean geometry called Riemannian geometry (also elliptic geometry) to Einstein, which was a necessary step in the development of Einstein's general theory of relativity. Abraham Pais's book on Einstein suggests that Grossmann mentored Einstein in tensor theory as well. Grossmann introduced Einstein to the absolute differential…

Grossmann collaborated with Einstein on GR, supplying quite a bit of the mathematical capacity required (which initially didn't come easily to Einstein). They published jointly, until Einstein was competent enough to work independently [1]. That's not equivalent to the situation being claimed here.

[1] https://arxiv.org/pdf/1312.4068

Re: More questions about whether researchers can trust OpenAI with unpublished math

#95
post #67
post #51

Earlier quoted context omitted.

It's hilarious how people think they care about their reputation, and wouldn't circumvent ZDR policies. Like bro, they literally covertly hired Apple employees and had them steal IP and equipment form Apple. They aren't scared of Apple lawyers, so they definitely aren't scared of yours.

what the point and usefulness of the comments above? we shouldn't be surprised? is normal to steal? hiring apple employees? can you realize what this means? focus on this part: "If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself" don't threat this as a minor d…

> we shouldn't be surprised? is normal to steal?

Two different things. It's not normal to steal, but we shouldn't be surprised thieves steal. It's what they do.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#96

Earlier quoted context omitted.

The difference is that Einstein didn't literally have someone prompting him towards his result.

Uh, he did. Marcel Grossmann. “It was Grossmann who emphasized the importance of a non-Euclidean geometry called Riemannian geometry (also elliptic geometry) to Einstein, which was a necessary step in the development of Einstein's general theory of relativity. Abraham Pais's book on Einstein suggests that Grossmann mentored Einstein in tensor theory as well. Grossmann introduced Einstein to the absolute differential…

Yeah and we get a nice list of attributions for who developed which idea, while OpenAI just takes credit for everything its model spits out.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#97

Earlier quoted context omitted.

Uh, he did. Marcel Grossmann. “It was Grossmann who emphasized the importance of a non-Euclidean geometry called Riemannian geometry (also elliptic geometry) to Einstein, which was a necessary step in the development of Einstein's general theory of relativity. Abraham Pais's book on Einstein suggests that Grossmann mentored Einstein in tensor theory as well. Grossmann introduced Einstein to the absolute differential…

Yeah and we get a nice list of attributions for who developed which idea, while OpenAI just takes credit for everything its model spits out.

Correction: OpenAI takes credit for what it's model spits out in response to other people's prompts. That's even worse.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#98
post #62

I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal. For a long time it was clearly the former, but now I think it is the latter. The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agent…

I think so too. The value is in the entire conversation. IMO, "domain experts" don't run LLMs blindly and hands free. This does not work for top level work (e.g., mathematical proofs, coding anything more complex than yet another slop game or website). Experts have long sessions where they prompt and guide LLM in response to what it produces. This is the discovery process. And frontier labs definitely train on that.

The billion dollar question is whether this works "out of the distribution". I.e., whether LLMs can only find and use the specific ideas buried in training data, or whether they can learn to apply the "thinking process" to a new problem. IMO this is still unanswered (due to these recent controversies).

But regardless of the answer, it seems we have a planet-scale positive feedback loop here. LLM became good (enough) by training on generally available data (books, internet, github) + RLFH, so experts tried to use them on hard tasks, which required lots of hand holding. These conversations became part of the training data, and the next generation of frontier LLMs were better. So, more experts used them on harder tasks, again requiring hand holding. These conversation became part of the training data... etc.

In a nutshell, top human minds across the world are pouring their skills into LLMs just by using them. This is not "continuous learning", but if you re-train on the most recent sessions every, say, quarter (which seems to be happening?) you get close to that in practice.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#99
If mathematician was already using OpenAI for research purpose and making progress due to inputs from OpenAI's responses, then I wouldn't put it beyond OpenAI's reach to generate different relevant prompts to make progress by itself. Afterall, Model can keep at it for whatever timeline and keep pursuing all possible combinations it can think try.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#100
post #5

Earlier quoted context omitted.

I would say there is a significant difference between AI discovering this completely on its own versus AI creating the finishing connecting part by connecting relevant data. Maybe this claim is too strong, but if part of it is true then the claims that OpenAI have made would be too strong as well. To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capt…

But this is what we do. Nobody ever invented or discovered anything in a vacuum - all discovery is synthesis of existing ideas and concepts applied to a novel domain. We laud Einstein for instance, but his work was a logical extension of Riemann - Riemann had a neat mathematical toy, Einstein described the universe with it - should we say Einstein was incapable of unique work?

[deleted]
Post reply on HN