Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

391–400 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#391

Earlier quoted context omitted.

I'm very happy to say calculators are far better than me in calculations (to a given precision). I'm happy to admit computers are so much better than me in so many aspects. And I have problem saying LLMs are very helpful tools able to generate output so much better than mine in almost every field of knowledge. Yet, whenever I ask it to do something novel or creative, it falls very short. But humans are ingenious beas…

But the question isn't whether you can get LLMs to do something novel, it's whether anyone can get them to do something novel. Apparently someone can, and the fact that you can't doesn't mean LLMs aren't good for that.

Novel is a tricky word. In this case, the LLM produced a python program that was similar to other programs in its corpus, and this oython program generated examples of hypergraphs that hadn't been seen before.

That's a new result, but I don't know about novel. The technique was the same as earlier work in this vein. And it seems like not much computational power was needed at all. (The article mentions that an undergrad left a laptop running overnight to produce one of the previous results, that's absolute peanuts when compared to most computational research).

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#392

Earlier quoted context omitted.

Our intelligence is related to brain structures, not all intelligence. You can't get to things like "what all intelligence, in general, is" from "what our intelligence is" any more than you can say that all food must necessarily be meat because sausages exist.

But... we're talking about our intelligence. So obviously it's quite relevant. I didn't say that AI isn't intelligent, I said that we have good reason to believe that our intelligence is unique. And we do, a lot of good evidence. I obviously don't believe that all intelligence is related to specific brain structure. Again, I'm a functionalist, so I believe that any structure that can exhibit the necessary functions w…

This is too dependent on what you mean by "unique", though. What do we have that apes don't, and which directly enables intelligence? What do we have that LLMs don't? What do LLMs have that we don't?

I don't think we know enough to definitively say "it's this bit that gives us intelligence, and there's no way to have intelligence without it". We just see what we have, and what animals lack, and we say "well it's probably some of these things maybe".

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#393

Earlier quoted context omitted.

What was the solution?

Well, I'm not going to share either solution as this is actually a pretty useful utility that I plan on releasing, but the short answer is: 1) don't use ScreenCaptureKit, and 2) take advantage of what CGWindowListCreateImage() offers through the content server. This is a simple IPC mechanism that does not trigger all the SKC limitations (i.e., no multi-space or multi-desktop support). In fact, when using SKC, the use…

Huh, Claude one-shotted it out of a single message from me. Man, LLMs have gotten good.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#394
post #378
post #350

Earlier quoted context omitted.

Do we know for a fact that LLMs aren't now configured to pass simple arithmetic like this in a simpler calculator, to add illusion of actual insight?

You can train a LLM on just multiplication and test it on ones it has never seen before, it's nothing particularly magical.

It's not 'magic' though but previously LLMs have performed very badly on longer multiplication, 'insight' is the wrong word but I'm saying maybe they're not wildly better at this calculation... maybe they are just optimising these well known jagged edges.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#395
post #391

Earlier quoted context omitted.

But the question isn't whether you can get LLMs to do something novel, it's whether anyone can get them to do something novel. Apparently someone can, and the fact that you can't doesn't mean LLMs aren't good for that.

Novel is a tricky word. In this case, the LLM produced a python program that was similar to other programs in its corpus, and this oython program generated examples of hypergraphs that hadn't been seen before. That's a new result, but I don't know about novel. The technique was the same as earlier work in this vein. And it seems like not much computational power was needed at all. (The article mentions that an underg…

I have never seen a human produce a Python program that wasn't similar to other programs they'd seem.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#396

Earlier quoted context omitted.

> I think there's demonstrably very little difference at all between human and AI outputs Bold claim, as the internet is awash with counterexamples. In any case, as I think this conversation is trending towards theories of artistic expression, “AI content” will never be truly relatable until it can feel pleasure, pain, and other human urges. The first thing I often think about when I critically assess a piece of art,…

> Bold claim, as the internet is awash with counterexamples. What do you consider a counterexample? Because I've been involved in local politics lately, and can say from experience that any foundation model is capable of more rational and detailed thought, and more creative expression, than most of the beloved members of my community. If you're comparing AI to the pinnacle of human achievement, as another commenter p…

>as another commenter pointed to Shakespeare

Lol wut?

I was not saying that LLMs cannot produce something like pinnacle of human achievement. I was saying we cannot quantify the difference between Shakespeare and something commonplace, because it requires the ability to feel.

I think you are being very dishonest here..

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#397

Earlier quoted context omitted.

I don't get what either of your points is intended to demonstrate. Let's revisit the first post I replied to: > It's deeply surprising to me that LLMs have had more success proving higher math theorems than making successful consumer software As far as I can tell, they absolutely have not had more success in this area relative to making successful consumer software.

Well we are kind of arguing past each other aren't we ? "More success" is a bit vague in this instance but building a compiler that would take a programmer 1 to 3 months is not comparable to this result regardless of whatever similarity exists in time completion estimates. That's the point. You can publish a paper (and in fact the researchers plan to) off this result. A basic compiler is cool but otherwise unremarkab…

> "More success" is a bit vague in this instance but building a compiler that would take a single programmer 1 to 3 months is not comparable to this result regardless of whatever similarity exists in time completion estimates. That's the point.

I guess we just disagree on this. It's not clear to me that these are totally different in terms of what they represent.

> You can publish a paper (and in fact the researchers plan to) off this result. A basic compiler is cool but otherwise unremarkable.

Publishing papers means very, very little to me. I can publish a paper on a programming language, you know that, right?

> You are leaning too hard on how long the researchers (who again did not manage to solve the problem in their attempts) estimated this would take and the "moderately interesting" tag of again, what was an open research problem.

I obviously estimate my "leanings" as being appropriate. I'm just using the researchers direct quotes. Factually, they had already come up with the approach that ultimately panned out. Factually, they estimated that a human could do this in some timeframe. What am I overly leaning on here?

> This, alongside a few math results that have cropped up in the last few months is easily more impressive than the vast majority of work being done with LLMs for software.

I think both are impressive, I don't know that I would draw some sort of big conclusions about it at this point. I definitely wouldn't draw the conclusion that AI is better at formal mathematics than producing software.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#398
post #285

Earlier quoted context omitted.

Humans are obviously unique in an interesting way. People only "move the goalpost" because it's not an interesting question that humans can do some great stuff, the interesting question is where the boundary is. (Whether against animals or AI). Some example goals which makes human trivially superior (in terms of intelligence): invention of nuclear bomb/plants, theory of relativity, etc.

But that's unique in the sense of "you have a bag of ten apples and I have a bag of eleven apples, therefore my bag is unique". It's not qualitatively different intelligence than a dog's, you just have more of it.

I would argue that point. The biological components are the same, but emergent behavior is a thing. So both the scale and the number of connections/way they connect have surpassed some limit after which cognitive capabilities increased severalfold to the point that humans "took over the world".

And arguably further increase in intelligence seems to fall into a diminishing returns category, compared to this previous boom. (Someone being "2x smarter" doesn't give them enough benefit of reigning over others, at least history would look otherwise were it the case, in my opinion)

Probably dumb example, but just by increasing speed you get well-behaving laminar flow vs turbulence, yet it's fundamentally the same a level beneath.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#399

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

> e.g. 167,383 * 426,397 = 71,371,609,051 They may be wrong, but so are you.

You could have just checked the math yourself, you know.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#400

Earlier quoted context omitted.

> I think there's demonstrably very little difference at all between human and AI outputs Bold claim, as the internet is awash with counterexamples. In any case, as I think this conversation is trending towards theories of artistic expression, “AI content” will never be truly relatable until it can feel pleasure, pain, and other human urges. The first thing I often think about when I critically assess a piece of art,…

> Bold claim, as the internet is awash with counterexamples. What do you consider a counterexample? Because I've been involved in local politics lately, and can say from experience that any foundation model is capable of more rational and detailed thought, and more creative expression, than most of the beloved members of my community. If you're comparing AI to the pinnacle of human achievement, as another commenter p…

The claim was precise:

> I think there's demonstrably very little difference at all between human and AI outputs

Counterexamples range from em-dashes, “Not-this, but-that”, people complaining about AI music on Spotify (including me) that sounds vaguely like a genre but is missing all of the instrumentation and motifs common to that genre.

The rest of your comment I don’t even know how to respond to, to be honest.

Post reply on HN