https://artificialanalysis.ai/models
Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?
501–510 of 1001 posts
https://artificialanalysis.ai/models
Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…
• 98.6% on ARC-AGI-3 • 97.6% on frontier math • 95.9% on CAD • 100% on ExploitBench Nothing modest about it
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
To each their own. Personally I will start feeling the AGI as soon as we move from chatting about benchmark results to learn that some lab just announced the discovery of tens of novel treatments for rare diseases. Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.
Secret tip to win the mario cart clone: Just hold w, no steering needed.
It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
> Like, what's the point, if the next AI can do it in 5 seconds? Live a life doing whatever makes you happy. Post-work society is an inevitability if we don't destroy our planet.
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…
Hey Astra, can you fix openai website so that static website doesn't lag on M3 Pro when I scroll?
great, but nobody can use it for another 100 days right?
“allowing non-technical people to create and play custom games that go beyond rudimentary elements” Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart
Here's a one-shotted submarine game I made with Fable a few weeks back - https://roryok.com/games/deepdive3d.html. One prompt, and I think it's deeper than this (if you'll pardon the pun)