Live data from Hacker News

GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

cdn.openai.com

441–450 of 467 posts

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#441

It seems like a solid set of criteria for how easily a task can be automated by AI agents is: - extent to which correctness of solution be easily specified and checked - extent to which new potential solutions can be implemented as text - extent to which prior art exists online This basically maps to software engineering and math. I think a fair bit of AI hype comes from the fact that the very architects of AI are th…

Interesting take! I feel like 2 of them are maybe overstated: > - extent to which correctness of solution be easily specified and checked I don't think most software is like solving a math problem or series of math problems. Algorithmic problems are very narrow and might be more like this though, where an oracle that verifies answers as either correct or incorrect exists beforehand. The correctness function of most s…

> I don't think most software is like solving a math problem or series of math problems.

I agree with you when talking about high level software design. As you say it ultimately boils down to building something people will pay for, which is a fuzzy correctness function that is hard to measure within an agentic sandbox.

But unlike other professions, there are a lot of sub-problems within software development that are able to be fully specified and tested via text generation. And I think the developers of AI overestimate how many such problems exist for other professions. What I’m saying is most other professions tend to be “fuzzy all the way down”… which incidentally is why they select for people with fuzzier skillsets. Or in other cases, like physical engineering, the correctness is quantitative, but the necessary I/O integrations and physical automation lower the ROI of agentic workflows considerably.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#442

Earlier quoted context omitted.

Something I've noticed is that if you run Qwen 3.6 35B-A3B (Q8) with a low temperature of 0.4, and leave default reasoning turned on, it will spend quite a lot of time in reasoning/thinking mode. But often it does figure out how to solve something on its own by correcting itself within its reasoning loop before it outputs the final 'answer'. If you watch the progress of the reasoning in llama-server while it's doing…

You'd have less problems with 27B, btw.

I didn't mean so much that it was a problem, but actually in some projects for exploring what's possible, watching its "thinking" mode output at temp 0.4 is intentional and useful to take notes and begin exploring new directions. Sometimes it'll come up with something I hadn't thought of, perhaps it will go down that direction, perhaps it'll disregard it...

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#443
post #28

Since this isn't in Lean and it's extremely easy for something like this to contain a subtle mistake, I think I'd prefer this be announced by a professional mathematician. The proof appears relatively short and elementary (not to be confused with easy -- just not using any advanced or modern machinery) so it shouldn't take long for the mathematics community to do a peer review. Without that, you could easily crank ou…

https://github.com/openai/cdc-lean

Perfect -- that's great to see. The proof strategy in Lean appears essentially identical to the natural language strategy (as much as is reasonably possible). I think this settles it!

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#444
post #318

Earlier quoted context omitted.

To be clear, I am not agreeing that people only upvote their favorite billionaires, just that if this particular thing was done by a Chinese model it would not have gotten the same attention

> a Chinese model it would not have gotten the same attention Well, you would be wrong.

Interesting point I guess, much to consider

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#445

If all checks out this is a huge milestone. AI has now solved one of the most famous open problems in graph theory, using an off the shelf model, in one hour. It might be a better mathematician than most humans at this point. Kind of like when chess software started beating everyone except grandmasters. What’s left? Proposing and building out entirely new theories and frameworks? Then better than any human? Then alie…

>> What’s left? For example, there's all the problems that the same off-the-shelf model hasn't solved despite OpenAI running it for many hours on them. Don't forget you're only seeing the results of successful runs. We can estimate that those unsolved problems must number in the dozens, or even hundreds, given the amount of time that passed since the last announcement of a solution to an interesting problem by an Ope…

The 2 most notable/interesting solutions have come from Open AI directly, but most of the 'LLM solves open problem' category didn't and has come from 3rd parties doing their own thing with publicly available models. I don't see why one would assume they're running models on hundreds of problems. Most likely they have a few problems they especially care about that they run on.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#446
post #253

Unrelated to the accomplishment or proof itself, but it's interesting how much of the prompt, even in this latest-and-greatest model, is spent essentially telling the model to actually solve the problem. Things like "Reject status reports, vague optimism, and claims that an unproved global compatibility statement is 'routine'." Also a lot prompt spent feeding it strategies, which feel like they should/will eventually…

Yes, the prompt, and use of subagents is interesting. It could be characterized as tree of thoughts rather than "think step by step" chain of thoughts. I see the need for this as coming down to two things: 1) LLMs are fundamentally prediction machines, and therefore ultimately will only do what they are prompted to do (and whatever that leads to). They may have been trained on, and/or have access to, all sorts of inf…

It might also be helpful more than 'necessary'. The 2 most notable solutions have come from open ai themselves, but most of the 'LLM solves open problem' category are from 3rd parties doing their own thing.

This one was pretty impressive in its own right (probably the most impressive outside these 2), and the prompt is concise and basic.

https://www.scientificamerican.com/article/amateur-armed-wit...

https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#447
post #253

Unrelated to the accomplishment or proof itself, but it's interesting how much of the prompt, even in this latest-and-greatest model, is spent essentially telling the model to actually solve the problem. Things like "Reject status reports, vague optimism, and claims that an unproved global compatibility statement is 'routine'." Also a lot prompt spent feeding it strategies, which feel like they should/will eventually…

Optimism and status reports burn extra tokens and make the user more prone to ask the model to process the problem again because it was "so close to solving".

This way you get more profit per API user and subscription users reach their quota faster and are contacted to update their plan to a higher tier.

It's working exactly as intended

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#448

Earlier quoted context omitted.

'roll down hill' is a good way of putting it. They don't have 'will', but that's as we want it I think. I think alignment is harder if they develop will. Without will they are still tools that feel like an exoskeleton rather than something that will control us.

Agents have a state which will unfold as a plan, especially in planning mode. Why not call this 'will'?

I think because without the initial prompt, they are only interpreting our will, and do not act under their own volition.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#449

Earlier quoted context omitted.

That is an incredibly insulting comment. I am married, two children, have lead a fantastic fulfilling life. Just because I don't believe the Flying Spaghetti Monster created the universe doesn't mean I am an "Edgelord". Remember the phrase, you also don't believe in God. There are hundreds of gods you don't believe in.

Bringing up the Flying Spaghetti Monster does make you an edgelord though.

Because it reminds you that your specific deity is just as ridiculous?

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#450

Earlier quoted context omitted.

It's hard for me not to think what's the point. I am a very average, even below average person in times of intelligence. What is even my value or reason to be if I know anything I can do, LLMs can do better? What is even my value both on job market and as a human?

There are smarter and better humans at just about everything you or I could want to do, that's just life. Most of life isn't about comparative advantages, it's about enjoying life with people we like.

I can't enjoy life with my people if I can't eat, drink, and have a roof over my head.
Post reply on HN