Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

351–360 of 850 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#351
post #104

This is the second wake up call. Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”. Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge…

This has been my line of thinking as well. I have developed a sort of paranoia when I'm working using AI on my projects. Who's to say Claude or OpenAI isn't using the final conclusion of all my ideas, trial and error, and adding it to their database of insights to be offered to the next subscriber for a price? They have demonstrated both the intelligence at scale and the lack of morals for this to not be a problem at…

[deleted]

Re: More questions about whether researchers can trust OpenAI with unpublished math

#352
post #104

This is the second wake up call. Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”. Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge…

This has been my line of thinking as well. I have developed a sort of paranoia when I'm working using AI on my projects. Who's to say Claude or OpenAI isn't using the final conclusion of all my ideas, trial and error, and adding it to their database of insights to be offered to the next subscriber for a price? They have demonstrated both the intelligence at scale and the lack of morals for this to not be a problem at…

Earlier in late 90s "to organize the world's information and make it universally accessible and useful." sounded cool. Now it has taken a sinister turn.

From being able to quickly find information and gain knowledge for the people, it is becoming - using information to manipulate and control the people.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#353

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

Both can be true: 1. OpenAI couldn't have solved the problem without the researchers' private data for training. 2. OpenAI models can solve math problems

Anthropic isnt getting enough scrutiny for their unprofessionalism:

1. Anthropic employee working on monumental problem but didnt receive/ask for the full backing of the company's resources

2. May or may not be mixing unreleased Claude output with Codex without zero data retention agreement

3. Victory lap on Twitter and giggling around the city before they finished the job, sparking rumors for competitors

Re: More questions about whether researchers can trust OpenAI with unpublished math

#355

I think this is mathematicians coming to grips with the fact that AI is surpassing them we will all have this moment soon enough, and it will change how we think about intelligence, identity and value

Or the companies hosting the frontier AIs are leeching the conversations.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#356
post #309

Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings. My naive instincts would be that it seems unlikely that a single chat transcript wou…

Yes. See https://www.anthropic.com/research/small-samples-poison?from... . 250 documents ingested from somewhere is enough to become part of the knowledge of a model of arbitrarily large size. I would expect that a good idea that fits in a framework that is already being ingested would be more easily taken up than some random thing unassociated with anything else. Could that go down to a single transcript? If the mod…

Thats not what they are asking. This paper is discussing documents in the training dataset poisoning the LLM for malicious behavior. This person are asking if anyone has deliberately put something in a private chat (presumably with retrain on my data turned off), to see if they can get it to leak across sessions from distinct users. I am positive this happens but I have not seen the proof. I also want to know the answer to this question.

Here are potentially relevant documents?

https://medium.com/secludy/fine-tuning-llm-on-sensitive-data...

https://spylab.ai/blog/non-adversarial-reproduction/

https://arxiv.org/abs/2601.18834

Re: More questions about whether researchers can trust OpenAI with unpublished math

#357

https://x.com/markchen90/status/2097400166554993041?s=20 that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.

Mark isn't saying the toggle does nothing.

He's saying that if you leave it on, your data can be used to help train our models.

If you opt out, we don't train on your data.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#358
post #293

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

If your rumor is true, what we are witnessing is a giant paradigm shift rather than individual incidents. Mathematicians were the first victims of super-intelligence. Of course it’s not an endless source. They had to burn millions of dollars to solve a single problem.

Wouldn't that be chess players as the first victims?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#359
post #293

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

If your rumor is true, what we are witnessing is a giant paradigm shift rather than individual incidents. Mathematicians were the first victims of super-intelligence. Of course it’s not an endless source. They had to burn millions of dollars to solve a single problem.

[deleted]

Re: More questions about whether researchers can trust OpenAI with unpublished math

#360

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

On your second point: there is a more plausible explanation which David Bessis calls the "overhang". The short version is that there is a large amount of relatively low hanging fruits in mathematics, because no human has broad enough knowledge and enough time to try them all. AI is not constraint by that, and therefore can systematically pluck all those low hanging fruits.

Quote: "The Overhang consists of the unrealized capital gains of past mathematical creativity, the latent value from connecting the dots in the existing corpus. It is a dividend of canonization. Mathematician X states problem A, mathematician Y crafts concept B, then mathematician Z notices that B trivially solves A and “captures” the social reward. But in the process of capturing the reward, Z usually introduces new concepts and new open problems, reinjecting latent value into the Overhang.

LLMs can be trained on the entirety of the mathematical corpus. Thanks to their phenomenal memorization and pattern-matching abilities (without always being able to map out their associative logic and attribute due credits), they are in a unique position to harvest the Overhang. By contrast, professional mathematicians have typically read a few hundred articles in their career, out of millions of existing references, less than 0.1% of the total.

This will lead to great discoveries, which is unambiguously exciting. But it could also lead to a sad new deal, where human slaves painfully curate the Overhang while AIs systematically beat them at the finish line."

source: https://substack.com/inbox/post/183753276

Post reply on HN