Earlier quoted context omitted.
Saying that the fundamental limitations are things like counting the number of rs in strawberry is boring, though. That's how tokens work and it's trivial to work around. Talking about how they find it hard to say they aren't sure of something is a much more interesting limitation to talk about, for example.
> Talking about how they find it hard to say they aren't sure of something is a much more interesting limitation to talk about, for example. Sure, thank you for steelmanning my argument. I didn’t think I needed to actually spell out all of the fundamental limitations of LLMs in this specific thread. They are spoken at length across the web, but are often met with pushback, which was my entire point. Here’s another on…
Epoch confirms GPT5.4 Pro solved a frontier math open problem
411–420 of 744 posts
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#412Earlier quoted context omitted.
>Publishing papers means very, very little to me. I can publish a paper on a programming language, you know that, right? We both know that you are not getting that published in a reputable journal without a lot of effort beyond merely 'publishing the language I created', but sure, I'm sure you can get something on arxiv. >I obviously estimate my "leanings" as being appropriate. I'm just using the researchers direct q…
> We both know that you are not getting that published in a reputable journal without a lot of effort beyond merely 'publishing the language I created'. But sure, you can get something on arxiv. lol what? There are papers on programming languages all the time. > 1. One is something that has been done many times before and the other an unsolved problem. It doesn't take a genius to see one estimate is likely much stron…
Sure and have you read them ? They're the results of many months or years of research and development so I really don't know what point you think you are making here.
>Building a compiler for a new programming language, building net new code, etc, is all stuff that was unsolved / had not been done before.
Okay but that's not taking a month or two or being asked of LLMs x10000 every day so thanks for making my point I guess.
>Feel free to explain the difference, I guess.
No thanks. If you don't understand it that's fine. This has run its course anyway.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#413Earlier quoted context omitted.
We already know exactly what causes these bugs. They are not a fundamental problem of LLMs, they are a problem of tokenizers. The actual model simply doesn't get to see the same text that you see. It can only infer this stuff from related info it was trained on. It's as if someone asked you how many 1s there are in the binary representation of this text. You'd also need to convert it first to think it through, or use…
Okay but, genuinely not an expert on the latest with LLMs, but isn’t tokenization an inherent part of LLM construction? Kind of like support vectors in SVMs, or nodes in neural networks? Once we remove tokenization from the equation, aren’t we no longer talking about LLMs?
Basically it's an engineering tradeoff. There is more demand for LLMs that can solve open math problems, but can't count the Rs in strawberry, than there is for models that can count letters but are bad at everything else.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#414Earlier quoted context omitted.
The claim was precise: > I think there's demonstrably very little difference at all between human and AI outputs Counterexamples range from em-dashes, “Not-this, but-that”, people complaining about AI music on Spotify (including me) that sounds vaguely like a genre but is missing all of the instrumentation and motifs common to that genre. The rest of your comment I don’t even know how to respond to, to be honest.
> em-dashes, “Not-this, but-that” I've literally seen humans accusing other humans of being AI here on hackernews for these. Q.E.D.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#415Earlier quoted context omitted.
No. That's wrong. LLMs don't output the highest probability taken: they do a random sampling.
This was obviously a simplification which holds for zero temperature. Obviously top-p-sampling will add some randomness but the probability of unexpected longer sequences goes asymptotically to zero pretty quickly.
A bog standard random number generator or even a flipping coin can produce novel output at will. That's a weird thing to fault LLMs for? Novelty is easy!
See also how genetic algorithms and re-inforcement learning constantly solve problems in novel and unexpected ways. Compare also antibiotics resistances in the real world.
You don't need smarts for novelty.
Where I see the problem is producing output that's both high quality _and_ novel. On command to solve the user's problem.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#416Earlier quoted context omitted.
> We both know that you are not getting that published in a reputable journal without a lot of effort beyond merely 'publishing the language I created'. But sure, you can get something on arxiv. lol what? There are papers on programming languages all the time. > 1. One is something that has been done many times before and the other an unsolved problem. It doesn't take a genius to see one estimate is likely much stron…
>lol what? There are papers on programming languages all the time. Sure and have you read them ? They're the results of many months or years of research and development so I really don't know what point you think you are making here. >Building a compiler for a new programming language, building net new code, etc, is all stuff that was unsolved / had not been done before. Okay but that's not taking a month or two or b…
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#417I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…
Most inventions are an interpolation of three existing ideas. These systems are very good at that.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#418Earlier quoted context omitted.
Huh, Claude one-shotted it out of a single message from me. Man, LLMs have gotten good.
No it didn't. Like I said... it may have gotten something that worked but there is no way Claude got it to work while supporting multi-spaces, multi-desktops, and using under 2% cpu utilization. My solution can display app window content even when those windows are minimized, which is not something the content server supports. My point was that Claude realized all the SKC problems and came up with a solution that 99%…
Maybe, but that's the magic of LLMs - they can now one-shot or few-shot (Ngood enough for a specific user. Like, not supporting multi-desktops is fine if one doesn't use them (and if that changes, few more prompts about this particular issue - now the user actually knows specifically what they need - should close the gap).
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#419Earlier quoted context omitted.
> em-dashes, “Not-this, but-that” I've literally seen humans accusing other humans of being AI here on hackernews for these. Q.E.D.
You’re really going to make the claim that there are no counterexamples of human and AI output being indistinguishable on the internet? At least make the counterclaim that “those are from old models, not the newest ones”, that’s more intellectually invigorating than the comment you just provided.
Is that a claim I've made? I don't see it anywhere. I think a lot of people think that because they can get the AI to generate something silly or obviously incorrect, that invalidates other output which is on-par with top-level humans. It does not. Every human holds silly misconceptions as well. Brain farts. Fat fingers. Great lists of cognitive biases and logical fallacies. We all make mistakes.
It seems to me that symbolic thinking necessitates the use of somewhat lossy abstractions in place of the real thing, primarily limited by the information which can be usefully stored in the brain compared to the informational complexity of the systems being symbolized. Which neatly explains one cognitive pathology that humans and LLMs share. I think there are most certainly others. And I think all the humans I know and all the LLMs I've interacted with exist on a multidimensional continuum of intelligence with significant overlap.
I hereby rebuff your crude and libelous mischaracterization of my assertion. How's that? :)
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#420I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…
I've been working on a utility that lets me "see through" app windows on macOS [1] (I was a dev on Apple's Xcode team and have a strong understanding of how to do this efficiently using private APIs). I wondered how Claude Code would approach the problem. I fully expected it to do something most human engineers would do: brute-force with ScreenCaptureKit. It almost instantly figured out that it didn't have to "see th…