Live data from Hacker News

GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

cdn.openai.com

161–170 of 467 posts

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#161
post #106
post #28

Since this isn't in Lean and it's extremely easy for something like this to contain a subtle mistake, I think I'd prefer this be announced by a professional mathematician. The proof appears relatively short and elementary (not to be confused with easy -- just not using any advanced or modern machinery) so it shouldn't take long for the mathematics community to do a peer review. Without that, you could easily crank ou…

…and thank God it's not Lean.

Why not both? Not sure why you're presenting this as one or the other.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#162

This is not a remark about AI, but there's something funny about mathematics in that every novel result is broadly perceived as a big deal. We attach basically zero value to writing a new program that hasn't existed before, or a piece of text that hasn't existed before. It's boring, or even a net negative, unless you can show that the result benefits the world in some way. We'd find it weird if OpenAI put out a relea…

> rejected any notion of utility. It would be fundamentally wrong for you to ask what's the value of solving the Erdős–Hajnal conjecture; the value is that it's solved. I disagree. Mathematicians care about the utility of a result. It is just that they regard mathematical understanding as a valid type of utility, and that can be arbitrarily far removed from practical utility. But a proof that doesn't help anyone unde…

It seems in mathematics that the utility of a problem is directly correlated with how difficult it is to solve, for some odd reason. If I defined some pointless construction and it turned out to be very difficult to prove, it would automatically over time become considered a "high utility" mathematics problem (again, for some odd reason).

Mathematics is largely just smart people working on pointless puzzles, and only by coincidence do these puzzles turn out to have practical applications (it cannot be predicted). Or I guess all the obviously practical problems in mathematics have already been solved -- we're now in a world where math is rarely the limiting factor for human progress (like it was, say, pre-calculus; was FFT the last significant unblock from math?).

It's such a waste of the best human minds. Or maybe the best human minds are actually doing something else, maybe we only notice the handful of Terence Taos, not the hundreds of people of equal brilliance who realized pure math is pointless and decided to pursue physics, rocketry, or quantitative finance.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#163

are the references real? how do you think it got access to those papers? were they somehow already in the training data, or a result of web searches, Google scholar, etc? None of them include a web URL but in text some are super specific ("[3, Sections 2.1 and 3.1]" and "[8, p. 367]"). The references go back to 1954 (Chronologically sorted: 1954, 1973, 1975, 1976, 1978, 1979, 1981, 1985, 1987 and 1994.) Since referen…

Yes, reference 10 jumped out at me as well. I thought personal correspondence references typically include one of the authors of the paper.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#164

I don't really like these articles, because they seem extremely hard to verify. OpenAI has published a lot of stuff in the past where, upon close inspection, what they're saying is technically true but a lot less interesting or impressive than the headline. Except by the time anyone looks into it, the hype has moved on. It seems like there's maybe a thousand people in the world that can even say if this is good or no…

1. A lot more than 1000, you're off by more than one order of magnitude. It's definitely beyond my level of graph theory knowledge (undergrad level) but looking at the paper, it's not using any crazy machinery, and it's less than 3 pages.

2. Those people will say whether it's a good proof or not. We have other examples of interesting proofs from AI, we're really beyond the point of arguing whether it can produce any interesting math (though it seems to do much better at combinatorics than anything else).

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#165
post #159
post #155

Earlier quoted context omitted.

I didn't say they have no value. Just limited value. A novel readable proof that expands the horizons of human insight is certainly more valuable than a megabyte sized trychnobezoar of machine generated predicates.

You are assuming that the latter, once autonomously discovered and verified at scale, could not simply be translated into the former, also perhaps autonomously at scale (or otherwise selectively as determined by human interest, taste, and relevance).

Well we're literally discussing a human readable machine generated proof here yet you don't seem happy with that.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#166

I don't really like these articles, because they seem extremely hard to verify. OpenAI has published a lot of stuff in the past where, upon close inspection, what they're saying is technically true but a lot less interesting or impressive than the headline. Except by the time anyone looks into it, the hype has moved on. It seems like there's maybe a thousand people in the world that can even say if this is good or no…

I think you may be overindexing on the criticisms here. OpenAI has absolutely done impressive work in math already, and the criticisms are almost always based on the article that they initially published, usually available here in the HN comments within a few hours at most. Headlines will be headlines and hype guys will be hype guys, but OpenAI and Anthropic aren't lying and their bots are doing impressive work.

This one is a well-known problem with a brief, approachable proof, and they published the prompt.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#167

Reading the prompt is very interesting. I always wonder how they make these long-running prompts and I guess they literally just tell it to "keep going". After working with LLMs day-in, day-out an SWE for months, I feel like this could be greatly improved with something like a state machine of progress and proper orchestration. Instead of spinning up a ton of subagents to follow different paths, whip up some Markdown…

I mean you can just ask them to do exactly that.

Especially with GPT (5.5), I've been having a lot of issues with it just repeatedly stalling out. I had to build a quota monitoring skill so that it'd keep plowing forward until either the task was finished (in some way) or the quota budget was exhausted.

I also had issues with the compaction. Codex seems to compact... weirdly, resulting in the agent becoming a newborn after each compaction event. Telling it to use a notes file is basically essential and self-evident.

Now that I mention, I should probably refine this skill to monitor the context window fill as well, to work around this.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#168
all easily varifyable tasks can now be solved with money. this is worth paying attention to. math proofs are verifyable -> math proofs are easy now. you can think of other such tasks: cybersecurity, AI R&D/RSI, killing people, 3d-printing helpful tools, maxxing-out human health, manipulation, self-driving cars, anything that can be checked

all jobs in the future will be those can not be easily verifiably done. if you need a team of people to decide if you have been productive, and those people cant be automated, you're in luck.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#169
post #15

Earlier quoted context omitted.

The prompt was released, but not the cost of the result.

Assuming all 64 subagents were running for a full hour (the tweet states just under an hour): Throughput Output tokens Output cost ---------------------------- ------------- ----------- 40 tok/s (5.5 low) ~9.2M ~$275 55 tok/s (5.5 base) ~12.7M ~$380 70 tok/s (5.5 high) ~16.1M ~$485 750 tok/s (Sol Fast, $75/M) ~172.8M ~$13,000 Claude estimates that tool use / input tokens might add 10-15% on top of that depending on e…

Sol fast isn't the Cerebras 750 tok/s version, it's just 1.5x speed at 2.5x price

I assume they didn't use the Cerebras version for this since it's probably very supply-constrained right now

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#170
post #149
post #58

Earlier quoted context omitted.

> I complained in the group chat, that our didactic materials, specifically tasked with providing motivation and concrete examples, did not contain a single application, of this most richly applied field. > I was promptly pilloried, and shunned. Heh. In my day I may have participated in the pillorying. I do think that there is value/merit in professors mentioning real world applications, where they exist . What they'…

Hear me out on this one: For a lot of math departments, that is exactly why they teach this. Education is rooted in application. We have entire careers that depend on certain aspects of mathematics, so most companies gatekeep that career by a degree. The degree requires the class. The student taking the class may not even be old enough to drink alcohol yet, and they can't possibly be expected to know of all the appli…

I think for many people (myself included) understanding mathematics is rooted in application because it helps bridge the divide between intuition and rote memorization. Without the application, IMO instructors are doing a disservice to their students and pedagogy of mathematics itself. They’re intentionally ignoring a significant fraction of the class, unless they’re teaching some esoteric grad level pure math.
Post reply on HN