Can’t wait for this stuff to have quality of life increases for the average person. So far all I see is that AI has made owning a computer more expensive, made some jobs redundant, increased spam and distrust with questionable authenticity of content and of course made some Americans very rich.
Ten advances in mathematics and theoretical computer science
321–330 of 1001 posts
Re: Ten advances in mathematics and theoretical computer science
#322The GitHub repo with the Lean formalizations just came out a couple of hours ago: https://github.com/openai/ten-proofs It also links to a paper written by an LLM where the model "reconstructs how the proof came together" based on the unpublished reasoning traces: https://cdn.openai.com/pdf/reasoning-walkthroughs.pdf I wish they'd publish the prompts though!
Exact prompts haven't mattered for about a year now.
Re: Ten advances in mathematics and theoretical computer science
#323Earlier quoted context omitted.
the fact that such things have to be explicitly in the prompt points to the fact that the underlying system is still far from where it needs to be (basically, lacks basic understanding what a proof is)
No, that's not what this is. This is a warning to the LLM that coming back with partial results is not good enough. Take a grad student with a perfectly good understanding of what a proof is. Their supervisor gives them a major problem to work on. Almost always, the problem is too hard, the student comes back with partial results, and student and the supervisor iterate from there. Now imagine that they have an unusua…
To a real mathematician you would not have to list those explicitly, he/she would have understood that implicitly from 'give me a full proof'. That listing makes sense to say only to somebody who pretends to be a mathematician, but has not true understanding of how the math works. The models are getting better and better in this pretension, but prompts like that reveal that it is still just a pretension, not a true understanding.
Re: Ten advances in mathematics and theoretical computer science
#324Earlier quoted context omitted.
The ones I'm familiar with are big breakthroughs, but they are both counterexamples. Examples have an advantage in that once you have the example in hand and a sketch of the proof (which they have provided), then an expert can probably work out the details themselves. The sofic groups question was the outstanding question about sofic groups. Almost everyone thought that non-sofic groups existed, and there were plausi…
Interesting. Do you have any more specific insights into where you feel AI was a big value-add to these problems? I don't want to be overly dismissive of AI, but I also feel that the AI hype engine frequently positions claims as being 'ground-breaking' when they are really just interesting incremental results. The general consensus of developers is that AI can only do the work of a strong 'junior'. Yet as soon as we…
If it works better here than for programming, then I would guess it's because you can give it a very precise prompt, so you either solve the problem or you don't. If you read the prompts people have shared for problems like this, then the instructions are basically "Solve this problem. Don't give up early. Don't solve a similar problem."
Re: Ten advances in mathematics and theoretical computer science
#325Re: Ten advances in mathematics and theoretical computer science
#326Pretty cool. The impact of AI is getting undeniable, there aren’t many positions left to move the goalposts to at this stage, next they’ll have to be outside the stadium entirely. The sooner people can be broken out of their denial about all this the better, and we can start actually taking it seriously.
Never understood all this talk about moving goalposts - you understand that's how science works, right? We improve, we learn, we recalibrate our expectations based on what we've learned. If we never "moved the goalposts", we'd be stuck scoring the same goals over and over.
Not long ago many folks were saying AI was the same as the crypto bubble. No real useful technology and only hype.
Re: Ten advances in mathematics and theoretical computer science
#327Earlier quoted context omitted.
Exact prompts haven't mattered for about a year now.
Care to elaborate? Curious about this. Is this because LLMs have been geared towards understanding user user intent behind a prompt rather than following the instructions exactly?
As long as you are not missing important information, how you word the prompt does not have any effect.
Re: Ten advances in mathematics and theoretical computer science
#328Earlier quoted context omitted.
This isn't really about delivering - it's more about helping to understand the shape of problems that AI can solve right now. If they took 1000 problems and threw the model at it and it solved these ten, is there something we learn about these ten problems and the kinds of things that current AI is good at? That's very different from picking ten problems _at random_ and solving all of them successfully, which would s…
That's totally disjointed from anything in this thread. The main accusation is that openai is cherrypicking math problems and we should be against these results. As if a mathematical proof stops being provably correct because it was cherry picked And frankly these "concerns" ignore reality. In any research phd course you're actively told to bite off something small and likely to be provable so that you can prove it (…
The post that started this sub-thread asked:
> 1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?
I think it's an extremely relevant question to ask, because it helps us better understand the current state of AI being able to handle math, for exactly the reasons I outlined. I was arguing against the idea this is just a reactionary anti-AI kind of question to ask. It's not! You can be very impressed by what AI is capable of in math (I am) and still think those are really interesting things for OpenAI to disclose (I do).
OpenAI specifically called out a $2000 per problem average, which implies something that's probably not true ("if you throw $2k at us we'll solve an open problem for you"). It would be cool to know what the actual number is.
Re: Ten advances in mathematics and theoretical computer science
#329Any advances in theoretical physics yet? Are there any fundamental obstacles? I would have thought not, but I haven't seen anything reported.
The fundamental obstacle is that we have no conceivable way to produce the energy levels to test the predictions that new theoretical physics would produce. We are like over a dozen orders of magnitude off.
Re: Ten advances in mathematics and theoretical computer science
#330At what point can we say AGI has been achieved? What is the test? AI is solving mathematical problems that humans have not been able to solve for decades. Is that not enough? Sam Altman has said "If superintelligence can't discover novel physics, I don't think it's a superintelligence." Is that the test? How far away are we from AI discovering novel physics? It seems within reach.
Just maths isn't really general enough for the G in AGI.