Live data from Hacker News

GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

cdn.openai.com

401–410 of 467 posts

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#401

Over on r/math, one objection to the proof has been raised (more discussion is needed to know if it is a problem): https://old.reddit.com/r/math/comments/1uszk3d/openai_claims...

No, that's not a problem at all. It just the notation that's a bit weird.

For example, if e is the a-edge (first edge) from the u side and v is the b-edge (second edge) from the v side then g_{u,e} = 0, g_{v,e} = a so d_e = 0 + f(x2) where f(x2) is the flow (from Kilpatrick and Jaeger's NZ8F) on the first edge next to v.

I checked the whole thing with some surface reformulations on my side and it looks right to me.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#402

No one here actually cares about Cycle Double Cover Conjecture. I can demonstrate this by pointing out that the only time this conjecture was ever mentioned on the website was 14 years ago in a submission[1] that linked to a (now retracted) proof paper. That story received exactly zero upvotes. No one cared enough to upvote it and no one cared enough to ever mention this conjecture again. [1] https://news.ycombinator…

No one here actually cares about folding laundry. I can demonstrate it by pointing out at the absence of posts on that subject.

...but when an affordable robot that folds laundry becomes available, people here pay attention.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#403

No one here actually cares about Cycle Double Cover Conjecture. I can demonstrate this by pointing out that the only time this conjecture was ever mentioned on the website was 14 years ago in a submission[1] that linked to a (now retracted) proof paper. That story received exactly zero upvotes. No one cared enough to upvote it and no one cared enough to ever mention this conjecture again. [1] https://news.ycombinator…

No one here actually cares about folding laundry. I can demonstrate it by pointing out at the absence of posts on that subject. ...but when an affordable robot that folds laundry becomes available, people here pay attention.

Oh, for laundry to be solved!

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#405
post #253

Unrelated to the accomplishment or proof itself, but it's interesting how much of the prompt, even in this latest-and-greatest model, is spent essentially telling the model to actually solve the problem. Things like "Reject status reports, vague optimism, and claims that an unproved global compatibility statement is 'routine'." Also a lot prompt spent feeding it strategies, which feel like they should/will eventually…

Yes, the prompt, and use of subagents is interesting. It could be characterized as tree of thoughts rather than "think step by step" chain of thoughts.

I see the need for this as coming down to two things:

1) LLMs are fundamentally prediction machines, and therefore ultimately will only do what they are prompted to do (and whatever that leads to). They may have been trained on, and/or have access to, all sorts of information that may be useful to solve a problem, but their predictive nature is to only use that information if explicitly prompted to, else it remains "dark" and inaccessible other than by luck. You're essentially having to tell the model "solve this problem using techniques A, B & C", otherwise techniques A, B & C will be off the radar unless the model already associates them to the problem.

2) The fundamental reason this sort of brute force tree-of-thoughts "explore all avenues" prompting is necessary, is because the model itself has no inherent curiosity to explore. Humans work differently. Our behavior is also prediction based, but we are also built for problem solving and continual exploration/learning via traits like curiosity (driven by prediction failure).

Problem solving via search can to some extent be prompted for, as here, or achieved via an external harness, but impasse resolution via curiosity, directed exploration and continual learning (if/when something new/unpredicted is encountered) is trickier. You can't usefully prompt a predictive model to "be curious" since that will only cause it to predict what a curious person would do, rather than the model being curious in reaction to the specific gaps in it's own knowledge.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#406

If all checks out this is a huge milestone. AI has now solved one of the most famous open problems in graph theory, using an off the shelf model, in one hour. It might be a better mathematician than most humans at this point. Kind of like when chess software started beating everyone except grandmasters. What’s left? Proposing and building out entirely new theories and frameworks? Then better than any human? Then alie…

>> What’s left?

For example, there's all the problems that the same off-the-shelf model hasn't solved despite OpenAI running it for many hours on them. Don't forget you're only seeing the results of successful runs.

We can estimate that those unsolved problems must number in the dozens, or even hundreds, given the amount of time that passed since the last announcement of a solution to an interesting problem by an OpenAI model: i.e. the unit distance problem which was announced solved in 20 May this year. That's a couple of months, yes? We can be fairly certain that OpenAI have been trying to solve other problems all this time, first because they are hell bent on demonstrating that their models can do maths and second because we just got another result, but it took that long. They were obviously not twiddling their thumbs all this time.

So if OpenAI are running their model on a single proble for eight hours at a time (according to the prompt they released) they could be easily have run a few hundred instances of their model on the same number of open problems 156 times for each instance (53 days since 20 May, with a model running in three eight-hour sessions per 24 hour day). I mean the only restriction is the cost they're willing to pay for the inference.

So yeah, there's a lot left to do still, don't worry.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#407

They post this and then say it's too dangerous to make open source. this is proof that in reality it's To protect their market position

Have OpenAI been on a crusade of "too dangerous to open source" recently? They've pivoted their more to messaging to "competitive reasons" recently, which I appreciate, because it's honest.

FWIW, Gemma4 31B is already quite a capable cyber/security model; do some post-training with RL gyms on it focused on cyber tasks and harnesses for a week or two, and on the specific domain of security/vuln-finding/pen-testing, you'll end up with an extremely capable frontier-cyber model that's entirely under your control at a shocking 31B.

Because securing your codebase, or securing your company's codebase is critical, and I consider it both an ethical and professional responsibility as a developer. It shouldn't depend on whether a classifier fires or not.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#408

Earlier quoted context omitted.

You believe this based off what?

Based on these people not being idiots or charlatans? Why wouldn't they verify it, knowing that any shenanigans would certainly come to light?

This is spot on. And yet, when I try to publish my papers on P = NP with the proof reducing to "trust me bro, I checked this out carefully" I get shot down by irrate reviewers who demand to see my work. Why can't those people just believe what I say?

Maybe if I was an AI?

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#410
Let me stand here on the Skeptic's Corner and be skeptical, so that the users who complain about skeptical comments have someone to direct their ire at. You're welcome.

Right, so, first, I haven't looked at the proof. Graph theory is not my subject and it would probably take me a few days to get my head around the whole thing. If OpenAI's LLM was used to prove an important graph theory result, then that's very good for them and graph theory.

However, I have to note that it's been 52 days since 20 May, the last date that OpenAI announced their previous mathematical result (a disproof of the unit distance conjecture).

What have OpenAI been doing all this time? I am willing to bet a good percentage of my money that they were trying, and failing, to produce the current result, or possibly something even juicier (one of the Millenium prize problems maybe?). They are hell bent on showing that their models are good for maths and science so they're very unlikely to have sat there twiddling their thumbs until they suddenly sprang into action and prompted their LLM once to generate just one proof. They must have been running the thing constantly, multiple instances of it, over that entire period.

Going by the instruction to run for eight hours before returning or giving up in their released prompt [1], that means they could have made at most 156 attempts to solve this problem, each of which failed except the last one [2].

So what happened to those other 156 attempts? Are we ever going to see them?

More importantly, who was it that selected the announced result? Who decided that this result is an actual proof? Until now, every proof generated by an LLM has been verified either by human mathematicians, or by human mathematicians x a proof assistant. What happened this time?

Obviously, any claims that this result were produced "autonomously" must be evaluated according to the answer to that last question. So far, LLMs have been incapable of distinguishing between a correct and an incorrect proof, which is also why they need to be run multiple times until they generate a correct one. If something has changed, it'd be interesting to know.

Finally, a magic eight ball that's correct one time out of 156 may be useful; or it may not. I honestly have no idea. I think time will tell.

__________________

[1] "Spend at least 8 hours on this before even thinking of returning or giving up"

https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98...

[2] That's 52 days from 20 May, times 3 for each eight-hour attempt in a 24-hour day.

But note well that the X post says that the solution was produced in "just under one hour" so that means the model didn't really stick to the prompt's time limit. Which means there may have been considerably more than 156 attempts that we'll probably never know of.

Or even considerably more if the model ignored the time limit going the other way.

Post reply on HN