Enabling two settings tripled our scores on the ARC-AGI-3 benchmark
1–6 of 6 posts
Re: Enabling two settings tripled our scores on the ARC-AGI-3 benchmark
#2On one hand I agree with ARC team that harnesses can incorporate some logic and overfit to win the tests. But also, the specific harness that OpenAI wants to use is not cherry picked. What then should be the rules? Same harness for all or company’s standard commercial harness? Technically even that could incorporate some hacks.
Re: Enabling two settings tripled our scores on the ARC-AGI-3 benchmark
#3> Agents do best when they remember what they’ve done
Sometimes I am baffled that sentences like this are used as headlines. I know I’ll sound dismissive… but: duh?!
Re: Enabling two settings tripled our scores on the ARC-AGI-3 benchmark
#4I think this blog post might be a "reply" of sorts to Anthropic's tweet from some days earlier?
> On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model.
Re: Enabling two settings tripled our scores on the ARC-AGI-3 benchmark
#5> Agents do best when they remember what they’ve done Sometimes I am baffled that sentences like this are used as headlines. I know I’ll sound dismissive… but: duh?!
They had to point out the obvious because ARC-AGI-3 wasn't designed with that in mind for some reason.
Re: Enabling two settings tripled our scores on the ARC-AGI-3 benchmark
#6So is this how Opus 5 ran it?