Live data from Hacker News

GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

cdn.openai.com

371–380 of 467 posts

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#371
post #294

Earlier quoted context omitted.

It’s the second most important problem in all of graph theory on this database of open math problems: https://www.openproblemgarden.org/category/graph_theory?sort... Obviously not an exact measurement but to give you some sense of the importance of the problem

This doesn't contradict anything he said though. People on HN care only because an LLM proved a very difficult conjecture, not because we are independently interested in this conjecture.

Only? Seems remarkable to me.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#372

It seems like a solid set of criteria for how easily a task can be automated by AI agents is: - extent to which correctness of solution be easily specified and checked - extent to which new potential solutions can be implemented as text - extent to which prior art exists online This basically maps to software engineering and math. I think a fair bit of AI hype comes from the fact that the very architects of AI are th…

- the extent to which solutions could be implemented as text: not sure about that. AlphaFold is basically a mechanical/geometrical/Chemical problem. There are other scientific transformer based models.

- the extent which solutions exist online - if you have a strong verification tool, you can generate examples, you can generate feedback, i think you could start with small/smaller prior art

- the extent which solutions could be specified and checked - if you have a lot of priort art, maybe llm's can find the good "patterns" and compare against them, and at least get close to a good results - but you'd still need human verification.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#373

Earlier quoted context omitted.

It's hard for me not to think what's the point. I am a very average, even below average person in times of intelligence. What is even my value or reason to be if I know anything I can do, LLMs can do better? What is even my value both on job market and as a human?

Manual labor

Military will always like warm bodies ;)

Sometimes it's easy for me to imagine a bad future like this; where most of the men are forced into military as the only org that still has use for wetware, and women are valued only for their ability to birth new people...

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#375
post #253

Unrelated to the accomplishment or proof itself, but it's interesting how much of the prompt, even in this latest-and-greatest model, is spent essentially telling the model to actually solve the problem. Things like "Reject status reports, vague optimism, and claims that an unproved global compatibility statement is 'routine'." Also a lot prompt spent feeding it strategies, which feel like they should/will eventually…

I think a lot of this has to do with the post-training these models normally get. They are designed to answer basic questions with straightforward and short summary answers. They have the capacity to reason deeply, but they are not biased towards that unless prompted. I think it's because LLMs as they are in 2026 are both highly capable but also parlor tricks. They are not sentient, you just set them up with the cont…

Something I've noticed is that if you run Qwen 3.6 35B-A3B (Q8) with a low temperature of 0.4, and leave default reasoning turned on, it will spend quite a lot of time in reasoning/thinking mode. But often it does figure out how to solve something on its own by correcting itself within its reasoning loop before it outputs the final 'answer'.

If you watch the progress of the reasoning in llama-server while it's doing the thinking, you can track its progress. Sometimes the dead ends it goes down or things that it considers and then disregards are themselves something useful to re-prompt it with later, and send it 'rolling downhill', to use the metaphor of another commenter here, in another direction towards the same effort.

Putting 3.6 35B-A3B into a state that lets it spend a lot of time in its reasoning mode before outputting an answer is probably not something that a web based SaaS LLM would tolerate, because it would frustrate many of the non technical end users who want a LLM to spit out an answer now.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#376
post #371

Earlier quoted context omitted.

This doesn't contradict anything he said though. People on HN care only because an LLM proved a very difficult conjecture, not because we are independently interested in this conjecture.

Only? Seems remarkable to me.

Yes, the fact that an LLM managed to prove the conjecture is remarkable, but presumably you don't find the conjecture itself interesting.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#377
post #9

It's really neat that the prompt was released! I'm curious how many unsolved problems are tried against frontier models when they come out. Are we trying every problems against every release? What is the solve success rate? Is there a sub-community within Mathematics that is coordinating this effort? How much untapped opportunity is there here?

Very good question I can only answer for one subset tracked by Terence Tao

https://github.com/teorth/erdosproblems

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#378
post #253

Unrelated to the accomplishment or proof itself, but it's interesting how much of the prompt, even in this latest-and-greatest model, is spent essentially telling the model to actually solve the problem. Things like "Reject status reports, vague optimism, and claims that an unproved global compatibility statement is 'routine'." Also a lot prompt spent feeding it strategies, which feel like they should/will eventually…

I think a lot of this has to do with the post-training these models normally get. They are designed to answer basic questions with straightforward and short summary answers. They have the capacity to reason deeply, but they are not biased towards that unless prompted. I think it's because LLMs as they are in 2026 are both highly capable but also parlor tricks. They are not sentient, you just set them up with the cont…

Even Fable hallucinates. I had it tracking down some very obscure Ancient Greek inscriptions and the response just made up a translation/context for one inscription after "looking it up." Now, it was still a very particular thing and I really had to get into the weeds to push it to that point, but who knows how many other gaps, near or far, it will happily skip over just for the sake of coherence. I think this is an issue more primarily with LLMs than sensory systems like Waymos or all the ML applied to industrial processes--that really only requires pattern recognition, often very impressive and subtle pattern recognition but its no different from an artist learning to tell the difference between Prussian blue and Navy blue or a Sommelier learning the fine distinctions between various regions of Bordeaux. Language has many more avenues and introduces inherent contradictions that do not always lend themselves to easy resolution. But there are no alternatives paths visible to the models, there is only ever the next word; stochastic, in the sense that the possibility space is open; deterministic, in the sense that the final response is always a necessary result of every token that came before it in their total sequence. Thus, any response is constantly in the work of erasing any possible alternative, slowly narrowing down what can be written. If contradictions in language necessarily involve interpretation, then the models will only ever choose one at a time, and for them, it will always be the right one. But anyone who understands the subtleties of language can tell you that when it comes to determining the truth of an indeterminate statement, there is never just one right answer; or, rather, the answer which is taken to be the "right" one depends on the possibility of its own reversal into falsehood, if any argument has to be made to justify it.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#379

Earlier quoted context omitted.

I did this when I taught my third grader calculus on the train, when she asked a question about the train accelerating faster sometimes than other times. She loved it, but I was just taking advantage of children's natural curiosity. Do you have some examples that the adult could instigate, rather than waiting for the child to express curiosity?

I've used filling a tank, balloon, or bucket (rate of flow, can be subdivided to teach limits, and use weird shapes for teaching area under curve and interpolation en route to integrals), or the classic throwing a ball back and forth and trying to describe the shape, the distance it flies, peak speed vs peak height, figuring out how hard you are "actually" throwing instantaneously. Honestly as soon as you start think…

Yes, exactly, identifying things that change is the key. Thank you!

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#380
post #307
post #227

Earlier quoted context omitted.

> then that's a problem with the original question - not the solution itself I think there's a good counterexample to this: Atiyah/MacDonald proove the Nullstellensatz ultimately by using some trick involving determinants. They give a very nice theoretical treatment of the content and context of the theorem. But the proof at one crucial point uses techniques that live conceptually outside of this context: While its p…

I was baffled enough by this comment to take my copy and look. The Nullstellensatz is an exercise late in the book long after Noetherian rings are introduced and they don't even do the Rabinowitz trick in the hints as they have enough theory to hit it the hard way. Determinants are nowhere to be found.

It's been a few years (~ 15) since I read it.

The determinant trick I'm referring to is used in prop 2.4 (in the Version I found via Google).

They use the determinant of the adjugate matrix if I remember correctly... But they just hit the reader with this without any motivation or even naming it.

Post reply on HN