Live data from Hacker News

Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

schema-harness.github.io

71–80 of 86 posts

Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

#71
post #65

To be clear, we’ll want to see how this performs against the hold-out set. If it holds up, though, it’s a big deal, and kind of in line with the vibes this year, which I’d typify as ‘harness matters’. Maybe we’d upgrade to ‘harness matters immensely’ if this can 100% ARC-AGI-3 on existing models (more in the 13% range without this harness). I’m pretty excited to see what sort of generalization we come to over the nex…

I think I'd typify it as "ARC-AGI doesn't matter" more than "harness matters". Or maybe "harness matters for some very specific tasks".

ARC-AGI is a bit of a joke at this point. The first version was supposed to be hard for neural nets because it had few examples and a super secret test set, with a different distribution than the public set no less. It took a while, but the original ARC ultimately fell to exactly the approach it was supposed to be protected from, i.e. big data memorisation, thanks to data augmentation techniques that increased the available training examples to the point that the private test set was eventually overcome.

ARC-AGI 2 went the same way because it was basically the same kind of dataset except this time with some attempt to further defend it against LLMs with restrictions on the compute budget. And now ARC-AGI 3 is saturated within ... what is it, weeks? since its release. The fact that it's the public set that's beaten doesn't matter, when the score is 99%. Systems that can score ~90% on the public sets of the previous ARC's can comfortably reach 70-80% on the corresponding private test sets, as far as my eyballing of results suggests.

It is time to accept that the whole idea of ARC is for the dustbin. It does not measure what it's supposed to measure -fluid intelligence, reasoning, whatever it is today. Its original premise, that a system could only beat ARC if it possessed human-like core knowledge systems (a-la Elizabeth Spelke's theory) has been comprehensively refuted: none of the systems that have ever performed well on any version of ARC has made any attempt to represent core knowledge systems in any way, shape or form.

Ultimately, if your machine intelligence (let alone AGI) test relies on tricks like only giving a few examples or keeping a secret test set to defend itself against the dominant approach to machine intelligence... then it's not a useful machine intelligence test. Or it just doesn't measure machine intelligence but... something else. Who knows what.

Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

#72
post #12

it looks like what they are doing is using a frontier model to write a simulator for a game and then solve using it. it's not as impressive as it looks. the goals of Arc-AGI-like constructs is to get an IQ-like figure using raw'ish 2D measurement 'games' in the hope that it would signify something meaningful. what this harness does is get the model to write a simulator first, it's measuring something entirely differe…

> State grounding turns raw observations into objects, variables, and relations that can be tracked. Mechanism discovery finds how that state changes under an action and writes the rule as an executable program The way I'm reading this isn't that they are writing a game simulator, but rather that they have two things they are evolving - a perceptual model of the game mapping from pixels to objects, and a behavioral m…

it's recursion

Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

#73

Earlier quoted context omitted.

> State grounding turns raw observations into objects, variables, and relations that can be tracked. Mechanism discovery finds how that state changes under an action and writes the rule as an executable program The way I'm reading this isn't that they are writing a game simulator, but rather that they have two things they are evolving - a perceptual model of the game mapping from pixels to objects, and a behavioral m…

it's recursion

No - it's not the harness being updated, it's the code representing the action predictions (& perceptual model).

Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

#74
post #65

Earlier quoted context omitted.

I think I'd typify it as "ARC-AGI doesn't matter" more than "harness matters". Or maybe "harness matters for some very specific tasks".

ARC-AGI is a bit of a joke at this point. The first version was supposed to be hard for neural nets because it had few examples and a super secret test set, with a different distribution than the public set no less. It took a while, but the original ARC ultimately fell to exactly the approach it was supposed to be protected from, i.e. big data memorisation, thanks to data augmentation techniques that increased the av…

>It took a while, but the original ARC ultimately fell to exactly the approach it was supposed to be protected from, i.e. big data memorisation

No it didn't. People tried big data memorization, and it didn't work. Base LLMs (even with millions of synthetic examples) never solved ARC-AGI-1.

It took a real algorithmic advancement - reasoning models - to solve it.

Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

#75

Earlier quoted context omitted.

ARC-AGI is a bit of a joke at this point. The first version was supposed to be hard for neural nets because it had few examples and a super secret test set, with a different distribution than the public set no less. It took a while, but the original ARC ultimately fell to exactly the approach it was supposed to be protected from, i.e. big data memorisation, thanks to data augmentation techniques that increased the av…

>It took a while, but the original ARC ultimately fell to exactly the approach it was supposed to be protected from, i.e. big data memorisation No it didn't. People tried big data memorization, and it didn't work. Base LLMs (even with millions of synthetic examples) never solved ARC-AGI-1. It took a real algorithmic advancement - reasoning models - to solve it.

DeepSeek V3.2 was tried without reasoning and it got 57% on ARC AGI 1. It's a 7 month model, so I'm pretty confident that base LLMs would be able to solve ARC AGI 1 without reasoning/CoT.

Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

#76
post #75

Earlier quoted context omitted.

>It took a while, but the original ARC ultimately fell to exactly the approach it was supposed to be protected from, i.e. big data memorisation No it didn't. People tried big data memorization, and it didn't work. Base LLMs (even with millions of synthetic examples) never solved ARC-AGI-1. It took a real algorithmic advancement - reasoning models - to solve it.

DeepSeek V3.2 was tried without reasoning and it got 57% on ARC AGI 1. It's a 7 month model, so I'm pretty confident that base LLMs would be able to solve ARC AGI 1 without reasoning/CoT.

No, you are misunderstanding the paper.

https://arxiv.org/abs/2607.06764

The base model got 15%. They built an elaborate looping harness that allows it to burn 100k tokens "thinking" about the problem, which got the 57%. This is just an alternative approach to reasoning.

Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

#77

Earlier quoted context omitted.

ARC-AGI is a bit of a joke at this point. The first version was supposed to be hard for neural nets because it had few examples and a super secret test set, with a different distribution than the public set no less. It took a while, but the original ARC ultimately fell to exactly the approach it was supposed to be protected from, i.e. big data memorisation, thanks to data augmentation techniques that increased the av…

>It took a while, but the original ARC ultimately fell to exactly the approach it was supposed to be protected from, i.e. big data memorisation No it didn't. People tried big data memorization, and it didn't work. Base LLMs (even with millions of synthetic examples) never solved ARC-AGI-1. It took a real algorithmic advancement - reasoning models - to solve it.

>> It took a real algorithmic advancement - reasoning models - to solve it.

You could only claim that if the numbers of parameters and training tokens remained constant while "reasoning" was added to the base models, which is not the case.

So if you look at the ARC-AGI-1 leaderboard (https://arcprize.org/leaderboard), you can clearly see that the bigger a model the better it performs, and that's for the "reasoning" models, e.g. looking at the graph, Claude Opus 4 is at ~30%, Opus 4.5 is between ~60% and ~80% and Claude 4.7 is at ~90% [1].

Not surprising: LLMs continue to improve in performance as long as more resources are spent to train them. "Algorithmic" advances would show the trend line going the other way, i.e. tokens and parameters decreasing steadily while performance either staying the same or improving.

If you've observed something like that I'll be happy to be corrected but I haven't.

____________________

[1] Incidentally, Opus 4.6 slightly outperforms 4.7 and 4.8 with their -alleged- reduced total parameter count. 4.6 is at 94.0% while 4.7 is at 93.5% and 4.8 at 92.5%.

Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

#78

Earlier quoted context omitted.

>It took a while, but the original ARC ultimately fell to exactly the approach it was supposed to be protected from, i.e. big data memorisation No it didn't. People tried big data memorization, and it didn't work. Base LLMs (even with millions of synthetic examples) never solved ARC-AGI-1. It took a real algorithmic advancement - reasoning models - to solve it.

>> It took a real algorithmic advancement - reasoning models - to solve it. You could only claim that if the numbers of parameters and training tokens remained constant while "reasoning" was added to the base models, which is not the case. So if you look at the ARC-AGI-1 leaderboard ( https://arcprize.org/leaderboard ), you can clearly see that the bigger a model the better it performs, and that's for the "reasoning"…

I don't agree with your definition.

The point of reasoning models is that some tasks fundamentally require a certain number of serial steps. Base models are limited to learning parallel algorithms because of their parallel training, and so struggle on inherently-serial tasks like solving logic puzzles.

The advancement from reasoning is that it allows LLMs to learn a broader class of algorithms.

Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

#79

Earlier quoted context omitted.

>> It took a real algorithmic advancement - reasoning models - to solve it. You could only claim that if the numbers of parameters and training tokens remained constant while "reasoning" was added to the base models, which is not the case. So if you look at the ARC-AGI-1 leaderboard ( https://arcprize.org/leaderboard ), you can clearly see that the bigger a model the better it performs, and that's for the "reasoning"…

I don't agree with your definition. The point of reasoning models is that some tasks fundamentally require a certain number of serial steps. Base models are limited to learning parallel algorithms because of their parallel training, and so struggle on inherently-serial tasks like solving logic puzzles. The advancement from reasoning is that it allows LLMs to learn a broader class of algorithms.

Sorry, which definition do you mean?

The problem with "reasoning" is that for the most part it happens outside of training e.g. as CoT. The base model knows what it knows, it knows what it's trained on, and it can't really go much farther than that.

Now, starting with o3, "reasoning" models are probably (who knows exactly) trained on traces of reasoning, either from automated systems or from human experts, but that doesn't mean they learn any kind of algorithm, just that they learn to reproduce the behaviour of "algorithms" (or of reasoning humans). That can improve performance on certain kinds of task (the ones in the training set) up to a point, but you're not going to make a tiny model perform like one ten times larger just by that.

On the other hand, I do think that LLMs can be made smaller without losing a commensurable amount of performance, like e.g. the Tiny Recursion Model (TRM) which did OK at ARC 1 (45% I think). But then you lose a lot of functionality also. Basically the larger models probably have more parameters than they really need and that's something the industry seems to have realised, but you still need huge parameter counts to reach top performance anyway.

And then there's the training tokens, which aren't getting any fewer.

Re: Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

#80

Earlier quoted context omitted.

I don't agree with your definition. The point of reasoning models is that some tasks fundamentally require a certain number of serial steps. Base models are limited to learning parallel algorithms because of their parallel training, and so struggle on inherently-serial tasks like solving logic puzzles. The advancement from reasoning is that it allows LLMs to learn a broader class of algorithms.

Sorry, which definition do you mean? The problem with "reasoning" is that for the most part it happens outside of training e.g. as CoT. The base model knows what it knows, it knows what it's trained on, and it can't really go much farther than that. Now, starting with o3, "reasoning" models are probably (who knows exactly) trained on traces of reasoning, either from automated systems or from human experts, but that d…

>but you're not going to make a tiny model perform like one ten times larger just by that.

Small reasoning models do indeed outperform base models that are 10x larger, at these logic/reasoning tasks that require serial computation.

They do not outperform at tasks that rely more on world knowledge or memorization.

In most cases the base model cannot complete logic tasks at all, or only for very small instances; it's reasoning or nothing.

> but that doesn't mean they learn any kind of algorithm

They do indeed learn algorithms and can step through them with CoT. This is what allows reasoning models to, e.g. reliably multiply large numbers by applying the grade-school multiplication algorithm.

Post reply on HN