More questions about whether researchers can trust OpenAI with unpublished math
381–390 of 847 posts
Re: More questions about whether researchers can trust OpenAI with unpublished math
#382Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…
On your second point: there is a more plausible explanation which David Bessis calls the "overhang". The short version is that there is a large amount of relatively low hanging fruits in mathematics, because no human has broad enough knowledge and enough time to try them all. AI is not constraint by that, and therefore can systematically pluck all those low hanging fruits. Quote: "The Overhang consists of the unreali…
Many problems are solvable, but require months of work, and thousands of pages of proof. So people do not even try to create or verify the proof. AI changes that, it can verify and perhaps even simplify it, to more digestible form.
Re: More questions about whether researchers can trust OpenAI with unpublished math
#383It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.
I think people are focusing on the training data issue too much. If the data was contaminated, I can still blame that on negligence. But, at least with the Navier-Stokes solution, it's clear [^1] that they learned that Alpöge and Buckmaster were getting close to a solution and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt. What…
What accident is it when the system is designed to function that way?
Re: More questions about whether researchers can trust OpenAI with unpublished math
#384Earlier quoted context omitted.
OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way…
This is literally "We have investigated ourselves and found no wrongdoing" Why should we trust them?
Re: More questions about whether researchers can trust OpenAI with unpublished math
#385Earlier quoted context omitted.
OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way…
This is literally "We have investigated ourselves and found no wrongdoing" Why should we trust them?
Re: More questions about whether researchers can trust OpenAI with unpublished math
#386Earlier quoted context omitted.
“ On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems.” - https://openai.com/index/navier-stokes-solution/ They do not explicitly admit to knowing about NS specifically, but are e…
So then they DIDN'T "learn the secret to cracking the problem". They simply knew that part of the problem was solved. Knowing a problem can be solved and knowing the solution are not the same thing.
Re: More questions about whether researchers can trust OpenAI with unpublished math
#387Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…
If your rumor is true, what we are witnessing is a giant paradigm shift rather than individual incidents. Mathematicians were the first victims of super-intelligence. Of course it’s not an endless source. They had to burn millions of dollars to solve a single problem.
Spinning it negatively like that doesn't do anybody good.
Were mathematicians the "victims" of calculators? of Matlab?
Were writers the ""vIcTiMs"" of word processors?? (apparently yes, according to old TV shows about computers during the 1980s, that you can see on YouTube)
> "tHiS iS nOt ThE sAmE" — Everyone every time.
No, just look it up. Look into old magazines and TV shows or newspaper articles from whenever a disruptive new technology came out.
Re: More questions about whether researchers can trust OpenAI with unpublished math
#388Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings. My naive instincts would be that it seems unlikely that a single chat transcript wou…
I run such tests since a long time at chorasimilarity open notebook. I always used guest non login accounts. As a mathematician I was able to check two plagiates (by humans) with even such primitive means. But I have to mention that some things irk me in this conversation about math or science and AI. First, I see lots of attribution and other related problems, with certain impact for the researcher proffesion. But I…
Re: More questions about whether researchers can trust OpenAI with unpublished math
#389Earlier quoted context omitted.
If your rumor is true, what we are witnessing is a giant paradigm shift rather than individual incidents. Mathematicians were the first victims of super-intelligence. Of course it’s not an endless source. They had to burn millions of dollars to solve a single problem.
Wouldn't that be chess players as the first victims?
Re: More questions about whether researchers can trust OpenAI with unpublished math
#390This is the second wake up call. Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”. Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge…
And people thought Experts Systems were bad.