Live data from Hacker News

Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel

starfleetmath.com

91–100 of 119 posts

Re: Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel

#91

Earlier quoted context omitted.

Post-money people with side interests are what built the current western civilization.

No, underpaid nerds have built modern civilization.

Yeah, what did I say?

Re: Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel

#92
I feel like I'm seeing a maths+AI change from "let's test the limits of LLMs by seeing if they can do useful math" to "LLMs can do useful math, now let's solve lots of problems!", or put a different way the goal has shifted from "interesting exercise for AI" to "making a big difference in math". Am I correct?

Are there practical applications of any these problems being solved? No judgement implied, I'm well aware that "no" only means "not yet".

Re: Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel

#93
post #4

Very interesting, on many levels: first, the raw additional compute / search harness is worth reading about; huge numbers of Lean 4 theorems, thousands of vCPUs available for spreading out search, embedding databases of proofs, all very interesting. Second, the proofs -- I understand the Lean 4 proofs to be refereed by Fable, and generated by Chat 5.6 Sol. Unlike the leaked proof of the Cycle Double Cover Conjecture…

This is great feedback (thank you for taking the time), & you especially bring up a fair point on the writeups needing to be more human readable. I'll work on that

I’ve been working on one problem for three weeks with fable if you want the repo

Re: Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel

#94
post #60

My mouth is agape at the fact that this project is basically what I have been working on non-stop for the last three weeks and just yesterday gotten to the point of evaluating; hats off... I only have one novel proof (non-Erdos) and 13 first-time formalizations thus far. I still like doing maths by pen and paper, but this is fun too.

When you say "working on" what is your actual contribution? Like, what should I imagine you do? For most people who tell AIs what to do and are proud of it, it's sadly mostly sitting around and staring at "thinking" output, and steering a bit, so I'm curious what the work looks like.

Valid question. If I were further along and had the time to succinctly write up all my contributions, I would just point you at my blog post. I’m generally a poor communicator, so here goes nothing.

I designed and stood up a sovereign inference/compute on my intranet. It uses a trust model that allows for a controller (me) to spin up untrusted inference/forge machines for Lean, Sage, or other runtimes. Untrusted sandbox workers integrate directly into my custom harness as first class “attachments.” This is the “orchestration” layer. It’s mostly on open weights, by design.

I haven’t yet started SFTing since my examples corpus isn’t quite where I’d like it.

I have solved and formalized a one non-Erdos conjecture. I have formalized several pieces of another subfield that does not exist in Mathlib yet.

As for what I am currently working on, I have an idea I want to build out about how we might think about sieving algebraic structures to generating new, insightful conjectures.

Using LLMs and distributed compute in this context is just a consequence of needing tools to help visualize or materialize things that I am otherwise bad at so I could keep doing the interesting things myself.

Re: Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel

#95
post #16

Earlier quoted context omitted.

The thing about math is we don't usually know what is pure fancy and what is civilization altering until far after the discovery. Once in a while it's a real targeted crack at something practical but most often it's collecting things which seem trial until you use them together and suddenly you have computers running LLMs. If it were really just about funding people who like math to have fun then it's easy to do fore…

What is their pay going to be justified by once computers start conjecturing and proving theorems on their own?

Is this question just for mathematicians in isolation or does it imply the same for most other jobs too? I think the answer for the former is we don't, mathematics would become a hobby the same as we don't hire people to be human calculators anymore because we have machines better st it. For the latter, it depends - some say UBI while the machines do the work, others say dystopia ruled by the machines or their few owners, yet others say anything in between.

Put differently: There is no natural law of the universe of why we pay people to do work. It just works well for us currently. If it stops making sense we do something else that works well for that new currently.

Re: Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel

#96
post #92

I feel like I'm seeing a maths+AI change from "let's test the limits of LLMs by seeing if they can do useful math" to "LLMs can do useful math, now let's solve lots of problems!", or put a different way the goal has shifted from "interesting exercise for AI" to "making a big difference in math". Am I correct? Are there practical applications of any these problems being solved? No judgement implied, I'm well aware tha…

I don't disagree with you zingar. I think Erdős problems are great for testing a system's capability on genuinely hard math, which has value as a benchmark in itself, and maybe as a stepping stone toward more real-world impact. Your sentiment is well put.

To answer your question directly: most Erdős problems don't have practical applications on their own; the value is the techniques and the machine-checked proofs they leave behind. But there's more real world value in solving some of the FrontierMath Open Problems or Millennium Problems. There's a Venn diagram of "hard problems" and "real world impact" for sure.

Re: Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel

#97
post #75
post #65

Earlier quoted context omitted.

There still seems to be a difference between useless pure math research and useless science or useless philosophy. Science, even useless science, still has a subject matter that is relevant to us independently of science, the real world. And philosophy studies concepts (like "knowledge") that occur in natural language and thought, and those concepts are relevant to us independently of philosophy. But pure math is ent…

"useless" pure math often turns out to have scientific applications. And surely you don't mean all pure math, so your argument ends up being circular --- useless math is useless. You happen to think that this math is useless, but you might be, or turn out to be, wrong. Also, you're misusing the term "self-referential". You seem to mean that it's a closed system ... but it's not, since mathematicians interact with it.…

> "useless" pure math often turns out to have scientific applications.

I don't think that's true for a reasonable interpretation of "often". I'm pretty sure the vast majority of pure math research is and remains useless.

> Also, you're misusing the term "self-referential". You seem to mean that it's a closed system ... but it's not, since mathematicians interact with it.

Then which better term do you propose? The point was that its subject matter lies within itself, which is very different from science and philosophy.

> so your argument ends up being circular --- useless math is useless.

It's not circular: You can replace "pure" with "useless" and the argument stays the same. I contrasted useless math with useless science and useless philosophy for a reason.

> Finally: so what?

I'll just point out that "so what" is not a counterargument. If you agree with my point but find it unimportant: that's fine with me.

> Solving Erdős problems seems at least as justifiable as finding the trillionth digit of pi or writing an Apple ][ emulator in Brainfuck or playing those video games that you demonize by calling them addictive.

The difference between achieving a new record for a video game, and pure math research, is that nobody is confused about what the former is: It's a game, or a sport. Speed runners or chess champions or pi digit calculators don't confuse themselves with noble researchers advancing the frontier of human knowledge.

Re: Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel

#98
post #92

I feel like I'm seeing a maths+AI change from "let's test the limits of LLMs by seeing if they can do useful math" to "LLMs can do useful math, now let's solve lots of problems!", or put a different way the goal has shifted from "interesting exercise for AI" to "making a big difference in math". Am I correct? Are there practical applications of any these problems being solved? No judgement implied, I'm well aware tha…

At some point the effort shifts from proof of concept to exploration of impact. Of course doing proving at large scale provides further feedback for improving the AIs. Visual reasoning, for example, may still need work.

The big impact will be with scaling, for example complete autoformalization of existing math, and automatic exploration for new conjectures, with emphasis on how interesting they are. Automatic conjecture generation goes way way back, to the days of Lenat's AM system. Modern AI should do a far better job.

Re: Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel

#100
post #55

Earlier quoted context omitted.

Can you explain what you're using the local compute for vs. the API-based frontier models? That was entirely unclear to me. Are you running tool calls that include inference with local fine tunes? And fast math packages? Controlled by the frontier model agents? Is there a way folks can contribute to this?

Thank you for the questions echelon. 1) As far as the AI models go, we used GPT 5.6 Sol, Fable 5, and Gemini-2-embeddings across the system 2) Yes, the agents are given bash tools that allows them to interact with the preinstalled mathematics packages/dependencies that are on the VMs 3) This was a setup as a relatively quick project without much thought for future contributions, I will spend some time thinking about…

This is so great!

I hate to bother with more questions, but I'm just so curious about this.

If you could roughly sketch out your agentic harness loop in a sentence or two, what does it look like? Which model(s) do the driving? How is progress measured?

What's your daily/monthly budget for this look like, if you don't mind my asking?

Post reply on HN