Earlier quoted context omitted.
Math and coding competition problems are easier to train because of strict rules and cheap verification. But once you go beyond that to less defined things such as code quality, where even humans have hard time putting down concrete axioms, they start to hallucinate more and become less useful. We are missing the value function that allowed AlphaGo to go from mid range player trained on human moves to superhuman by p…
> I don't see this getting better. We went from 2 + 7 = 11 to "solved a frontier math problem" in 3 years, yet people don't think this will improve?
Epoch confirms GPT5.4 Pro solved a frontier math open problem
221–230 of 744 posts
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#222Earlier quoted context omitted.
I entered the prompt: > Write me a stanza in the style of "The Raven" about Dick Cheney on a first date with Queen Elizabeth I facilitated by a Time Travel Machine invented by Lin-Manuel Miranda It outputted a group of characters that I can virtually guarantee you it has never seen before on its own
Yes, but it has seen The Raven, it has seen texts about Dick Cheney, first dates, Queen Elizabeth, time machines and Lin Manuel Miranda. All of its output is based on those things it has seen.
Virtually all output from people is based in things the person has experienced.
People aren't designed to objectively track each and every event or observation they come across. Thus it's harder to verify. But we only output what has been inputted to us before.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#223Earlier quoted context omitted.
But human researchers are also remixers. Copying something I commented below: > Speaking as a researcher, the line between new ideas and existing knowledge is very blurry and maybe doesn't even exist. The vast majority of research papers get new results by combining existing ideas in novel ways. This process can lead to genuinely new ideas, because the results of a good project teach you unexpected things.
>But human researchers are also remixers. Some human researchers are also remixers to Some degree. Can you imagine AI coming up with refraction & separation lie Newton did?
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#224I have long said I am an AI doubter until AI could print out the answers to hard problems or ones requiring tons of innovation. Assuming this is verified to be correct (not by AI) then I just became a believer. I would like to see a few more AI inventions to know for sure, but wow, it really is a new and exciting world. I really hope we use this intelligence resource to make the world better.
most issues at every scale of community and time are political, how do you imagine AI will make that better, not worse? there's no math answer to whether a piece of land in your neighborhood should be apartments, a parking lot or a homeless shelter; whether home prices should go up or down; how much to pay for a new life saving treatment for a child; how much your country should compel fossil fuel emissions even when…
When I wrote that I hope we use it for good things, I was just putting a hopeful thought out there, not necessarily trying to make realistic predictions. It's more than likely people will do bad things with AI. But it's actually not set in stone yet, it's not guaranteed that it has to go one way. I'm hopeful it works out.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#225Earlier quoted context omitted.
yes its ridiculously good at stuff like that now. I dare you to try and trick it.
https://news.ycombinator.com/item?id=47495568
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#226Earlier quoted context omitted.
Yes, but is it "intelligence" is a valid question. We have known for a long time that computers are a lot faster than humans. Get a dumb person who works fast enough and eventually they'll spit out enough good work to surpass a smart person of average speed. It remains to be seen whether this is genuinely intelligence or an infinite monkeys at infinite typewriters situation. And I'm not sure why this specific example…
Someone actually mathed out infinite monkeys at infinite typewriters, and it turns out, it is a great example of how misleading probabilities are when dealing with infinity: "Even if every proton in the observable universe (which is estimated at roughly 1080) were a monkey with a typewriter, typing from the Big Bang until the end of the universe (when protons might no longer exist), they would still need a far greate…
What? Did you see one crying?
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#227Really? He was steering the wheel the whole time. GPT didn't do the math.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#228Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#229Earlier quoted context omitted.
Math and coding competition problems are easier to train because of strict rules and cheap verification. But once you go beyond that to less defined things such as code quality, where even humans have hard time putting down concrete axioms, they start to hallucinate more and become less useful. We are missing the value function that allowed AlphaGo to go from mid range player trained on human moves to superhuman by p…
Maybe to get a real breakthrough we have to make programming languages / tools better suited for LLM strengths not fuss so much about making it write code we like. What we need is correct code not nice looking code.
The bitter lesson is that the best languages / tools are the ones for which the most quality training data exists, and that's pretty much necessarily the same languages / tools most commonly used by humans.
> Correct code not nice looking code
"Nice looking" is subjective, but simple, clear, readable code is just as important as ever for projects to be long-term successful. Arguably even more so. The aphorism about code being read much more often than it's written applies to LLMs "reading" code as well. They can go over the complexity cliff very fast. Just look at OpenClaw.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#230Earlier quoted context omitted.
> The capabilities of a human far surpass every single AI to date Meaning however you (reasonably) define intelligence, if you compare humans to any AI system humans are overwhelmingly more capable. Defining "intelligence" as "solving a math equation" is not a reasonable definition of intelligence. Or else we'd be talking about how my calculator is intelligent. Of course computers can compute faster than we can, that…
>Meaning however you (reasonably) define intelligence, if you compare humans to any AI system humans are overwhelmingly more capable. Really ? Every Human ? Are you sure ? because I certainly wouldn't ask just any human for the things I use these models for, and I use them for a lot of things. So, to me the idea that all humans are 'overwhelmingly more capable' is blatantly false. >Defining "intelligence" as "solving…
Here might be some definitions of intelligence for example:
> The aggregate or global capacity of the individual to act purposefully, to think rationally, and to deal effectively with his environment.
> "...the resultant of the process of acquiring, storing in memory, retrieving, combining, comparing, and using in new contexts information and conceptual skills".
> Goal-directed adaptive behavior.
> a system's ability to correctly interpret external data, to learn from such data, and to use those learnings to achieve specific goals and tasks through flexible adaptation
But even a housefly possesses levels of intelligence regarding flight and spacial awareness that dominates any LLM. Would it be fair to say a fly is more intelligent than an LLM? It certainly is along a narrow set of axis.
> Because the only brute-forced aspect of LLM intelligence is its creation.
I would consider statistical reasoning systems that can simulate aspects of human thought to be a form of brute force. Not quite an exhaustive search, but massively compressed experience + pattern matching.
But regardless, even if both forms of intelligence arrived via some form of brute force, what is more important to me is the result of that - how does the process of employing our intelligence look.
> This very post, with the transcript available is an example of how untrue it is.
The transcript lacks the vector embeddings of the model's reasoning. It's literally just a summary from the model - not even that really.
> Do you realize how much compute it would take to run a full simulation of the human brain on a computer ? The most powerful super computer on the planet could not run this in real time.
You're so close to getting it lol