Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

1–10 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#2
> Subsequent to this solve, we finished developing our general scaffold for testing models on FrontierMath: Open Problems. In this scaffold, several other models were able to solve the problem as well: Opus 4.6 (max), Gemini 3.1 Pro, and GPT-5.4 (xhigh).

Interesting. Whats that “scaffold”? A sort of unit test framework for proofs?

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#3
post #2

> Subsequent to this solve, we finished developing our general scaffold for testing models on FrontierMath: Open Problems. In this scaffold, several other models were able to solve the problem as well: Opus 4.6 (max), Gemini 3.1 Pro, and GPT-5.4 (xhigh). Interesting. Whats that “scaffold”? A sort of unit test framework for proofs?

I think in this context, scaffolds are generally the harness that surrounds the actual model. For example, any tools, ways to lay out tasks, or auto-critiquing methods.

I think there's quite a bit of variance in model performance depending on the scaffold so comparisons are always a bit murky.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#8
post #2

> Subsequent to this solve, we finished developing our general scaffold for testing models on FrontierMath: Open Problems. In this scaffold, several other models were able to solve the problem as well: Opus 4.6 (max), Gemini 3.1 Pro, and GPT-5.4 (xhigh). Interesting. Whats that “scaffold”? A sort of unit test framework for proofs?

I think in this context, scaffolds are generally the harness that surrounds the actual model. For example, any tools, ways to lay out tasks, or auto-critiquing methods. I think there's quite a bit of variance in model performance depending on the scaffold so comparisons are always a bit murky.

Usually involves a lot of agents and their custom contexts or system prompts.
Post reply on HN