Neat. Maybe even deeply interesting. Absolutely garbage write up.
Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
41–50 of 86 posts
Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
#42In the spirit of ARC-AGI-3-like challenges, we just tested if frontier AI models are able to solve a lovely puzzle game, Baba Is You: https://quesma.com/blog/baba-is-bench/ A year ago, Sonnet 4 barely solved the first level. Now, both Fable 5 and GPT-5.6 Sol beat the first two stages. GPT 5.2 is slow, but efficient, while Gemini 3.1 Pro and 3.5 Flash struggle.
FWIW: "Baba Is You" is 7 years old and heralded as one of the greatest puzzle games of all times, with guides and solutions shared all over the internet. How to beat this game is 100% in the training set.
In a few instances (we covered it in Caveats) Gemini 3.5 Flash "knew" which level it was, but misremembered, and went with a wrong solution.
Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
#43it looks like what they are doing is using a frontier model to write a simulator for a game and then solve using it. it's not as impressive as it looks. the goals of Arc-AGI-like constructs is to get an IQ-like figure using raw'ish 2D measurement 'games' in the hope that it would signify something meaningful. what this harness does is get the model to write a simulator first, it's measuring something entirely differe…
Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
#44Any custom harness for a problem shows that harness engineering is going away. Eventually models will introspect problems, then build custom harnesses tailored to that. Then use and modify the ephemeral harness as required. Sol Ultra style is the path forward. The models are smart enough to self serve their tooling and processes. Given a problem they can figure it out and ask for directions when needed.
That’s called brute force so not really.
Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
#45Any custom harness for a problem shows that harness engineering is going away. Eventually models will introspect problems, then build custom harnesses tailored to that. Then use and modify the ephemeral harness as required. Sol Ultra style is the path forward. The models are smart enough to self serve their tooling and processes. Given a problem they can figure it out and ask for directions when needed.
I seriously doubt this, especially in a world in which there's not just one model. This makes sense if the models some how become unified.
Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
#46Any custom harness for a problem shows that harness engineering is going away. Eventually models will introspect problems, then build custom harnesses tailored to that. Then use and modify the ephemeral harness as required. Sol Ultra style is the path forward. The models are smart enough to self serve their tooling and processes. Given a problem they can figure it out and ask for directions when needed.
Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
#47In the spirit of ARC-AGI-3-like challenges, we just tested if frontier AI models are able to solve a lovely puzzle game, Baba Is You: https://quesma.com/blog/baba-is-bench/ A year ago, Sonnet 4 barely solved the first level. Now, both Fable 5 and GPT-5.6 Sol beat the first two stages. GPT 5.2 is slow, but efficient, while Gemini 3.1 Pro and 3.5 Flash struggle.
I'm wondering what's up with the release of Gemini 3.5 Pro, they keep postponing it. For a while, Google was doing pretty well with their releases.
Works fast - tells people how to overthrow the government.
Follows all rules and conventions Google wants - says corporate speak without actually accomplishing anything.
Can actually do complicated things- apt to tell the user to fuck off and do the hard work themselves.
Training models seems more akin to raising a kid than a computer application.
Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
#48it looks like what they are doing is using a frontier model to write a simulator for a game and then solve using it. it's not as impressive as it looks. the goals of Arc-AGI-like constructs is to get an IQ-like figure using raw'ish 2D measurement 'games' in the hope that it would signify something meaningful. what this harness does is get the model to write a simulator first, it's measuring something entirely differe…
on the flip side, the idea that most tests are bad, even standardized tests, the tests that you scored well on that gave you all your opportunities in life: it cuts to the emotional, grounded core, the absolute foundation, of too many people. in the crowd of hacker news commenters; people who buy anthropic shares at retail; the people who work at tech companies; and their kids, families, etc., who are a bunch of nobodies, there are a lot more incentives to believe "stupid fucking arcade games test AGI" than not.
Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
#49In the spirit of ARC-AGI-3-like challenges, we just tested if frontier AI models are able to solve a lovely puzzle game, Baba Is You: https://quesma.com/blog/baba-is-bench/ A year ago, Sonnet 4 barely solved the first level. Now, both Fable 5 and GPT-5.6 Sol beat the first two stages. GPT 5.2 is slow, but efficient, while Gemini 3.1 Pro and 3.5 Flash struggle.
FWIW: "Baba Is You" is 7 years old and heralded as one of the greatest puzzle games of all times, with guides and solutions shared all over the internet. How to beat this game is 100% in the training set.
Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
#50Earlier quoted context omitted.
The simulator the model builds is comparable to the mental model of the game humans create. It is also much more efficient, GPT 5.6 Sol cost $25,000 to run on ARC-AGI-3
> simulator the model builds is comparable to the mental model of the game humans create then they should try to use that for a more complicated game than Arc AGI. Arc games are simple by design, if you have the model simulate them they become trivial.
Eh, this is kind of sounds like being a prey animal that develops an almost unbeatable colored camouflage and then the predator develops infrared vision making it useless and the prey saying "no fair, you cheated".
People use algorithmic models all the time on problems that are far too difficult or large for their minds to conceptualize, is this not just an extension of that?