Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

401–410 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#401

Earlier quoted context omitted.

Well, I'm not going to share either solution as this is actually a pretty useful utility that I plan on releasing, but the short answer is: 1) don't use ScreenCaptureKit, and 2) take advantage of what CGWindowListCreateImage() offers through the content server. This is a simple IPC mechanism that does not trigger all the SKC limitations (i.e., no multi-space or multi-desktop support). In fact, when using SKC, the use…

Huh, Claude one-shotted it out of a single message from me. Man, LLMs have gotten good.

No it didn't. Like I said... it may have gotten something that worked but there is no way Claude got it to work while supporting multi-spaces, multi-desktops, and using under 2% cpu utilization. My solution can display app window content even when those windows are minimized, which is not something the content server supports.

My point was that Claude realized all the SKC problems and came up with a solution that 99% of macOS devs wouldn't even know existed.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#402
post #398

Earlier quoted context omitted.

But that's unique in the sense of "you have a bag of ten apples and I have a bag of eleven apples, therefore my bag is unique". It's not qualitatively different intelligence than a dog's, you just have more of it.

I would argue that point. The biological components are the same, but emergent behavior is a thing. So both the scale and the number of connections/way they connect have surpassed some limit after which cognitive capabilities increased severalfold to the point that humans "took over the world". And arguably further increase in intelligence seems to fall into a diminishing returns category, compared to this previous b…

Yeah, I don't know that there's such a jump. Dogs, for example, clearly communicate, both with us and with each other. They don't have language, but they also don't lack communication skills. To me, language is just "better communication" rather than a qualitatively different thing.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#403

Earlier quoted context omitted.

I'm very happy to say calculators are far better than me in calculations (to a given precision). I'm happy to admit computers are so much better than me in so many aspects. And I have problem saying LLMs are very helpful tools able to generate output so much better than mine in almost every field of knowledge. Yet, whenever I ask it to do something novel or creative, it falls very short. But humans are ingenious beas…

But the question isn't whether you can get LLMs to do something novel, it's whether anyone can get them to do something novel. Apparently someone can, and the fact that you can't doesn't mean LLMs aren't good for that.

To have a proper discussion we would have to define the word "novel" and that's a challenge in itself. In any case, millions of poeple tried to ask LLMs to do something creative and the results were bland. Hence my conclusion LLMs aren't good for that. But I'm also open they can be an element of a longer chain that could demonstrate some creativity - we'll see.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#404

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

LLMs can generate anything by design. LLMs can't understand what they are generating so it may be true, it may be wrong, it may be novel or it may be known thing. It doesn't discern between them, just looks for the best statistical fit.

The core of the issue lies in our human language and our human assumptions. We humans have implicitly assigned phrases "truly novel" and "solving unsolved math problem" a certain meaning in our heads. Some of us at least, think that truly novel means something truly novel and important, something significant. Like, I don't know, finding a high temperature superconductor formula or creating a new drug etc. Something which involver real intelligent thinking and not randomizing possible solutions until one lands. But formally there can be a truly novel way to pack the most computer cables in a drawer, or truly novel way to tie shoelaces, or indeed a truly novel way to solve some arbitrary math equation with an enormous numbers. Which a formally novel things, but we really never needed any of that and so relegated these "issues" to a deepest backlog possible. Utilizing LLMs we can scour for the solutions to many such problems, but they are not that impressive in the first place.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#405

Earlier quoted context omitted.

But... we're talking about our intelligence. So obviously it's quite relevant. I didn't say that AI isn't intelligent, I said that we have good reason to believe that our intelligence is unique. And we do, a lot of good evidence. I obviously don't believe that all intelligence is related to specific brain structure. Again, I'm a functionalist, so I believe that any structure that can exhibit the necessary functions w…

This is too dependent on what you mean by "unique", though. What do we have that apes don't, and which directly enables intelligence? What do we have that LLMs don't? What do LLMs have that we don't? I don't think we know enough to definitively say "it's this bit that gives us intelligence, and there's no way to have intelligence without it". We just see what we have, and what animals lack, and we say "well it's prob…

> What do we have that apes don't, and which directly enables intelligence?

Again, there are multiple fields of study with tons of amazingly detailed answers to this. We know about specific proteins, specific brain structures, we know about specific cognitive capabilities in the abstract, etc.

> What do we have that LLMs don't?

Again, quite a lot is already known about this.

This feels a bit like you're starting to explore this area and you're realizing that intelligence is complex, but you may not realize that others have already been doing this work and we have a litany of information on the topic. There are big open questions, of course, but we're definitely past the point of being able to say "there is a difference between human and ape intelligence" etc.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#406
post #396

Earlier quoted context omitted.

> Bold claim, as the internet is awash with counterexamples. What do you consider a counterexample? Because I've been involved in local politics lately, and can say from experience that any foundation model is capable of more rational and detailed thought, and more creative expression, than most of the beloved members of my community. If you're comparing AI to the pinnacle of human achievement, as another commenter p…

>as another commenter pointed to Shakespeare Lol wut? I was not saying that LLMs cannot produce something like pinnacle of human achievement. I was saying we cannot quantify the difference between Shakespeare and something commonplace, because it requires the ability to feel. I think you are being very dishonest here..

[dead]

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#407

Earlier quoted context omitted.

Well we are kind of arguing past each other aren't we ? "More success" is a bit vague in this instance but building a compiler that would take a programmer 1 to 3 months is not comparable to this result regardless of whatever similarity exists in time completion estimates. That's the point. You can publish a paper (and in fact the researchers plan to) off this result. A basic compiler is cool but otherwise unremarkab…

> "More success" is a bit vague in this instance but building a compiler that would take a single programmer 1 to 3 months is not comparable to this result regardless of whatever similarity exists in time completion estimates. That's the point. I guess we just disagree on this. It's not clear to me that these are totally different in terms of what they represent. > You can publish a paper (and in fact the researchers…

>Publishing papers means very, very little to me. I can publish a paper on a programming language, you know that, right?

We both know that you are not getting that published in a reputable journal without a lot of effort beyond merely 'publishing the language I created', but sure, I'm sure you can get something on arxiv.

>I obviously estimate my "leanings" as being appropriate. I'm just using the researchers direct quotes. Factually, they had already come up with the approach that ultimately panned out.

This really should not be hard to understand.

1. One is something that has been done many times before and the other an unsolved problem. It doesn't take a genius to see one estimate is likely much stronger than the other. If your point hinges on comparing them directly, it's pretty weak.

2. A moderately interesting open research problem is not the same thing as a moderately interesting problem and you seem to be conflating the two.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#408

Earlier quoted context omitted.

> Bold claim, as the internet is awash with counterexamples. What do you consider a counterexample? Because I've been involved in local politics lately, and can say from experience that any foundation model is capable of more rational and detailed thought, and more creative expression, than most of the beloved members of my community. If you're comparing AI to the pinnacle of human achievement, as another commenter p…

The claim was precise: > I think there's demonstrably very little difference at all between human and AI outputs Counterexamples range from em-dashes, “Not-this, but-that”, people complaining about AI music on Spotify (including me) that sounds vaguely like a genre but is missing all of the instrumentation and motifs common to that genre. The rest of your comment I don’t even know how to respond to, to be honest.

> em-dashes, “Not-this, but-that”

I've literally seen humans accusing other humans of being AI here on hackernews for these. Q.E.D.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#409

Earlier quoted context omitted.

Maybe because society has invested $trillions into this hammer and influencers are trying to convince CEOs to fire everyone and buy a bunch of hammers instead. My comment even said “LLMs have utility”. I gave an inch, and now the mile must be taken.

Saying that the fundamental limitations are things like counting the number of rs in strawberry is boring, though. That's how tokens work and it's trivial to work around. Talking about how they find it hard to say they aren't sure of something is a much more interesting limitation to talk about, for example.

> Talking about how they find it hard to say they aren't sure of something is a much more interesting limitation to talk about, for example.

Sure, thank you for steelmanning my argument. I didn’t think I needed to actually spell out all of the fundamental limitations of LLMs in this specific thread. They are spoken at length across the web, but are often met with pushback, which was my entire point.

Here’s another one: LLMs do not have a memory property. Shut off the power and turn it back on and you lose all context. Any “memory” feature implemented by companies that sell LLM wrappers are a hack on top of how LLMs work, like seeding a context window before letting the user interact with the LLM.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#410

Earlier quoted context omitted.

> "More success" is a bit vague in this instance but building a compiler that would take a single programmer 1 to 3 months is not comparable to this result regardless of whatever similarity exists in time completion estimates. That's the point. I guess we just disagree on this. It's not clear to me that these are totally different in terms of what they represent. > You can publish a paper (and in fact the researchers…

>Publishing papers means very, very little to me. I can publish a paper on a programming language, you know that, right? We both know that you are not getting that published in a reputable journal without a lot of effort beyond merely 'publishing the language I created', but sure, I'm sure you can get something on arxiv. >I obviously estimate my "leanings" as being appropriate. I'm just using the researchers direct q…

> We both know that you are not getting that published in a reputable journal without a lot of effort beyond merely 'publishing the language I created'. But sure, you can get something on arxiv.

lol what? There are papers on programming languages all the time.

> 1. One is something that has been done many times before and the other an unsolved problem. It doesn't take a genius to see one estimate is likely much stronger than the other.

Building a compiler for a new programming language, building net new code, etc, is all stuff that was unsolved / had not been done before.

> 2. A moderately interesting open research problem is not the same thing as a moderately interesting problem and you seem to be conflating the two.

Feel free to explain the difference, I guess.

Post reply on HN