Live data from Hacker News

GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

cdn.openai.com

171–180 of 467 posts

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#171
post #115

Earlier quoted context omitted.

Many harnesses include a current date and time in their system prompt, and if there is a way for the model to call for an updated time (either a dedicated time tool or calling the OS' `date` tool) they can track time they spent doing something. If not told up-front, they can try to infer it from timestamps in their logs. Sort of like a human - if you ask them to time something and give them a stopwatch, they do it. I…

Once on a late-night session, I had Cline!Claude spontaneously point out the time to me and suggest that I get to bed and come back fresh the next day. I don't think it's in the system prompt, but that the harnesses time-stamp each turn in the context. And from what I've seen, they also include the current and max context, so that the model can decide whether to continue work, suggest compaction, or prefer actions th…

> Once on a late-night session, I had Cline!Claude spontaneously point out the time to me and suggest that I get to bed and come back fresh the next day.

I had Claude say something "It's getting late, let's pick this up tomorrow" at like 11am.

As for context, in my experience Claude starts trying either to do maximum work with minimum tokens when it's approaching limit, or it starts deferring useful work while doing busy work. Both result in a mess and complete loss of traction after compaction.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#172

I don't really like these articles, because they seem extremely hard to verify. OpenAI has published a lot of stuff in the past where, upon close inspection, what they're saying is technically true but a lot less interesting or impressive than the headline. Except by the time anyone looks into it, the hype has moved on. It seems like there's maybe a thousand people in the world that can even say if this is good or no…

1. A lot more than 1000, you're off by more than one order of magnitude. It's definitely beyond my level of graph theory knowledge (undergrad level) but looking at the paper, it's not using any crazy machinery, and it's less than 3 pages. 2. Those people will say whether it's a good proof or not. We have other examples of interesting proofs from AI, we're really beyond the point of arguing whether it can produce any…

Right, but my criticism is to the hit-and-run nature of these hype pieces. By the time there's any semblance of what it actually means everyone has moved on but then you have a bunch of people operating under delusions from the hype. I get why OpenAI does it but I wish people would stop upvoting it. Like, hacker news is not a mathematics forum so the only purpose of this kind of thing is hype boosting or polarizing people. I am not looking forward to the "MATH IS SOLVED!" people for the next few days.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#173

Earlier quoted context omitted.

You say those things like they're a short step away, but that might not be how it works out. For example, AI has made zero progress in the last few years in surpassing professionals at art or writing. Its prompt-following skill is much better, and sure, it can render hands and text now, but its artistic sensibility is completely stagnant.

The difference is that artistic sensibility is largely subjective. This means that: 1. It's hard to measure (and people can disagree about it) 2. It can't really be improved using RL without a human in the loop (which is how math is being trained)

[deleted]

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#174

Earlier quoted context omitted.

> rejected any notion of utility. It would be fundamentally wrong for you to ask what's the value of solving the Erdős–Hajnal conjecture; the value is that it's solved. I disagree. Mathematicians care about the utility of a result. It is just that they regard mathematical understanding as a valid type of utility, and that can be arbitrarily far removed from practical utility. But a proof that doesn't help anyone unde…

It seems in mathematics that the utility of a problem is directly correlated with how difficult it is to solve, for some odd reason. If I defined some pointless construction and it turned out to be very difficult to prove, it would automatically over time become considered a "high utility" mathematics problem (again, for some odd reason). Mathematics is largely just smart people working on pointless puzzles, and only…

Writing that mathematics is a waste is such a hilariously ignorant comment to make on a programming forum.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#175
post #2

Announcement: https://x.com/__eknight__/status/2075643450196971805 Prompt: https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98...

> Spend at least 8 hours on this before even thinking of returning or giving up. Do current model harnesses have concepts of amount of time spent? Sometimes the model notices if a subprocess takes too long/hangs and kills it, but I've never seen it time itself.

that can run date

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#176

This is not a remark about AI, but there's something funny about mathematics in that every novel result is broadly perceived as a big deal. We attach basically zero value to writing a new program that hasn't existed before, or a piece of text that hasn't existed before. It's boring, or even a net negative, unless you can show that the result benefits the world in some way. We'd find it weird if OpenAI put out a relea…

> rejected any notion of utility. It would be fundamentally wrong for you to ask what's the value of solving the Erdős–Hajnal conjecture; the value is that it's solved. I disagree. Mathematicians care about the utility of a result. It is just that they regard mathematical understanding as a valid type of utility, and that can be arbitrarily far removed from practical utility. But a proof that doesn't help anyone unde…

[deleted]

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#177
post #15

Earlier quoted context omitted.

The prompt was released, but not the cost of the result.

Assuming all 64 subagents were running for a full hour (the tweet states just under an hour): Throughput Output tokens Output cost ---------------------------- ------------- ----------- 40 tok/s (5.5 low) ~9.2M ~$275 55 tok/s (5.5 base) ~12.7M ~$380 70 tok/s (5.5 high) ~16.1M ~$485 750 tok/s (Sol Fast, $75/M) ~172.8M ~$13,000 Claude estimates that tool use / input tokens might add 10-15% on top of that depending on e…

[deleted]

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#178
post #93

If all checks out this is a huge milestone. AI has now solved one of the most famous open problems in graph theory, using an off the shelf model, in one hour. It might be a better mathematician than most humans at this point. Kind of like when chess software started beating everyone except grandmasters. What’s left? Proposing and building out entirely new theories and frameworks? Then better than any human? Then alie…

> What's left? I think humans will be left to propose new conjectures while machines fill out the proofs. I don't know if there are enough interesting conjectures to go round to build new careers, though.

Surely the machines will have superior conjectures soon.

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#179

If all checks out this is a huge milestone. AI has now solved one of the most famous open problems in graph theory, using an off the shelf model, in one hour. It might be a better mathematician than most humans at this point. Kind of like when chess software started beating everyone except grandmasters. What’s left? Proposing and building out entirely new theories and frameworks? Then better than any human? Then alie…

You say those things like they're a short step away, but that might not be how it works out. For example, AI has made zero progress in the last few years in surpassing professionals at art or writing. Its prompt-following skill is much better, and sure, it can render hands and text now, but its artistic sensibility is completely stagnant.

[deleted]

Re: GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

#180
post #121

This is not a remark about AI, but there's something funny about mathematics in that every novel result is broadly perceived as a big deal. We attach basically zero value to writing a new program that hasn't existed before, or a piece of text that hasn't existed before. It's boring, or even a net negative, unless you can show that the result benefits the world in some way. We'd find it weird if OpenAI put out a relea…

Isn't it immediately obvious that solving something that humans have been unable to do for decades or more is the most tangible proof of ASI, or at the very least pretty good AGI?

Is this something humans have been unable to do?

There’s only so many people with the necessary skills to solve this. And you need these humans to choose to spend their time solving this, and not something else.

Post reply on HN