Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

411–420 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#411

Earlier quoted context omitted.

Saying that the fundamental limitations are things like counting the number of rs in strawberry is boring, though. That's how tokens work and it's trivial to work around. Talking about how they find it hard to say they aren't sure of something is a much more interesting limitation to talk about, for example.

> Talking about how they find it hard to say they aren't sure of something is a much more interesting limitation to talk about, for example. Sure, thank you for steelmanning my argument. I didn’t think I needed to actually spell out all of the fundamental limitations of LLMs in this specific thread. They are spoken at length across the web, but are often met with pushback, which was my entire point. Here’s another on…

But that's also like saying "humans don't have a memory property, any 'memory' is in the hippocampus". It's not useful to say that "an LLM you don't bother to keep training has no memory". Of course it doesn't, you removed its ability to form new memories!

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#412

Earlier quoted context omitted.

>Publishing papers means very, very little to me. I can publish a paper on a programming language, you know that, right? We both know that you are not getting that published in a reputable journal without a lot of effort beyond merely 'publishing the language I created', but sure, I'm sure you can get something on arxiv. >I obviously estimate my "leanings" as being appropriate. I'm just using the researchers direct q…

> We both know that you are not getting that published in a reputable journal without a lot of effort beyond merely 'publishing the language I created'. But sure, you can get something on arxiv. lol what? There are papers on programming languages all the time. > 1. One is something that has been done many times before and the other an unsolved problem. It doesn't take a genius to see one estimate is likely much stron…

>lol what? There are papers on programming languages all the time.

Sure and have you read them ? They're the results of many months or years of research and development so I really don't know what point you think you are making here.

>Building a compiler for a new programming language, building net new code, etc, is all stuff that was unsolved / had not been done before.

Okay but that's not taking a month or two or being asked of LLMs x10000 every day so thanks for making my point I guess.

>Feel free to explain the difference, I guess.

No thanks. If you don't understand it that's fine. This has run its course anyway.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#413

Earlier quoted context omitted.

We already know exactly what causes these bugs. They are not a fundamental problem of LLMs, they are a problem of tokenizers. The actual model simply doesn't get to see the same text that you see. It can only infer this stuff from related info it was trained on. It's as if someone asked you how many 1s there are in the binary representation of this text. You'd also need to convert it first to think it through, or use…

Okay but, genuinely not an expert on the latest with LLMs, but isn’t tokenization an inherent part of LLM construction? Kind of like support vectors in SVMs, or nodes in neural networks? Once we remove tokenization from the equation, aren’t we no longer talking about LLMs?

It's not a side effect of tokenization per se, but of the tokenizers people use in actual practice. If somebody really wanted an LLM that can flawlessly count letters in words, they could train one with a naive tokenizer (like just ascii characters). But the resulting model would be very bad (for its size) at language or reasoning tasks.

Basically it's an engineering tradeoff. There is more demand for LLMs that can solve open math problems, but can't count the Rs in strawberry, than there is for models that can count letters but are bad at everything else.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#414

Earlier quoted context omitted.

The claim was precise: > I think there's demonstrably very little difference at all between human and AI outputs Counterexamples range from em-dashes, “Not-this, but-that”, people complaining about AI music on Spotify (including me) that sounds vaguely like a genre but is missing all of the instrumentation and motifs common to that genre. The rest of your comment I don’t even know how to respond to, to be honest.

> em-dashes, “Not-this, but-that” I've literally seen humans accusing other humans of being AI here on hackernews for these. Q.E.D.

You’re really going to make the claim that there are no counterexamples of human and AI output being indistinguishable on the internet? At least make the counterclaim that “those are from old models, not the newest ones”, that’s more intellectually invigorating than the comment you just provided.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#415
post #115
post #105

Earlier quoted context omitted.

No. That's wrong. LLMs don't output the highest probability taken: they do a random sampling.

This was obviously a simplification which holds for zero temperature. Obviously top-p-sampling will add some randomness but the probability of unexpected longer sequences goes asymptotically to zero pretty quickly.

I'm not sure what the point is?

A bog standard random number generator or even a flipping coin can produce novel output at will. That's a weird thing to fault LLMs for? Novelty is easy!

See also how genetic algorithms and re-inforcement learning constantly solve problems in novel and unexpected ways. Compare also antibiotics resistances in the real world.

You don't need smarts for novelty.

Where I see the problem is producing output that's both high quality _and_ novel. On command to solve the user's problem.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#416

Earlier quoted context omitted.

> We both know that you are not getting that published in a reputable journal without a lot of effort beyond merely 'publishing the language I created'. But sure, you can get something on arxiv. lol what? There are papers on programming languages all the time. > 1. One is something that has been done many times before and the other an unsolved problem. It doesn't take a genius to see one estimate is likely much stron…

>lol what? There are papers on programming languages all the time. Sure and have you read them ? They're the results of many months or years of research and development so I really don't know what point you think you are making here. >Building a compiler for a new programming language, building net new code, etc, is all stuff that was unsolved / had not been done before. Okay but that's not taking a month or two or b…

Yeah idk what you're going on about lol

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#417

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

Most inventions are an interpolation of three existing ideas. These systems are very good at that.

My take as well. Furthermore, most innovations come relatively shortly after their technological prerequisites have been met, so that suggests the "novelty space" that humans generally explore is a relatively narrow band around the current frontier. Just as humans can search through this space, so too should machines be capable of it. It's not an infinitely unbounded search which humans are guided through by some manner of mystic soul or other supernatural forces.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#418

Earlier quoted context omitted.

Huh, Claude one-shotted it out of a single message from me. Man, LLMs have gotten good.

No it didn't. Like I said... it may have gotten something that worked but there is no way Claude got it to work while supporting multi-spaces, multi-desktops, and using under 2% cpu utilization. My solution can display app window content even when those windows are minimized, which is not something the content server supports. My point was that Claude realized all the SKC problems and came up with a solution that 99%…

> it may have gotten something that worked but there is no way Claude got it to work while supporting multi-spaces, multi-desktops, and using under 2% cpu utilization.

Maybe, but that's the magic of LLMs - they can now one-shot or few-shot (Ngood enough for a specific user. Like, not supporting multi-desktops is fine if one doesn't use them (and if that changes, few more prompts about this particular issue - now the user actually knows specifically what they need - should close the gap).

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#419

Earlier quoted context omitted.

> em-dashes, “Not-this, but-that” I've literally seen humans accusing other humans of being AI here on hackernews for these. Q.E.D.

You’re really going to make the claim that there are no counterexamples of human and AI output being indistinguishable on the internet? At least make the counterclaim that “those are from old models, not the newest ones”, that’s more intellectually invigorating than the comment you just provided.

> claim that there are no counterexamples of human and AI output being indistinguishable on the internet?

Is that a claim I've made? I don't see it anywhere. I think a lot of people think that because they can get the AI to generate something silly or obviously incorrect, that invalidates other output which is on-par with top-level humans. It does not. Every human holds silly misconceptions as well. Brain farts. Fat fingers. Great lists of cognitive biases and logical fallacies. We all make mistakes.

It seems to me that symbolic thinking necessitates the use of somewhat lossy abstractions in place of the real thing, primarily limited by the information which can be usefully stored in the brain compared to the informational complexity of the systems being symbolized. Which neatly explains one cognitive pathology that humans and LLMs share. I think there are most certainly others. And I think all the humans I know and all the LLMs I've interacted with exist on a multidimensional continuum of intelligence with significant overlap.

I hereby rebuff your crude and libelous mischaracterization of my assertion. How's that? :)

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#420

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

I've been working on a utility that lets me "see through" app windows on macOS [1] (I was a dev on Apple's Xcode team and have a strong understanding of how to do this efficiently using private APIs). I wondered how Claude Code would approach the problem. I fully expected it to do something most human engineers would do: brute-force with ScreenCaptureKit. It almost instantly figured out that it didn't have to "see th…

Why is ScreenCaptureKit a bad choice for performance?
Post reply on HN