The first question I still have is what happened to core knowledge priors. The white paper that introduced ARC made a big todo about how core knowledge priors are necessary to solve ARC tasks but from what I can tell none of the best-performing (or at-all performing) systems have anything to do with core knowlege priors. So what happened to that assumption? Is it dead? The second question I still have is about the de…
Even the strongest possible interpretation of the results wouldn't conclude "ARC-AGI is dead" because none of the submissions came especially close to human-level performance; the criteria was 85% success but the best in 2024 was 55%. That said, I think there should be consideration via information thermodynamics: even with TTT these program-generating systems are using an enormous amount of bits compared to a human…
This isn’t my area of expertise, but it seems plausible to me that what you said is completely erroneous or at the very least completely unverifiable at this point in time. How do you quantify how many bits it takes a human mind to solve one of the ARC problems?
That seems likely beyond the level of insight we have into the structure of cognition and information storage etc etc in wetware. I could of course be wrong and would love to be corrected if so! You mentioned a “tiny portion” of the human mind, but (as far as I’m aware), any given “small” part of human cognition still involves huge amounts of complexity and compute.
Maybe you are saying that the high level decision making a human goes through when solving can be represented with a relatively small number of pieces of information/logical operations (as opposed to a much lower level notion closer to the wetware of the quantity of information) but then it seems unfair to compare to the low level equivalent (weights & biases, FLOPs etc) in the ML system when there may be higher order equivalents.
I do appreciate the general notion of wanting to normalize against something though, and some notion of information seems like a reasonable choice, but practically out of our reach. Maybe something like peak power or total energy consumption would be a more reasonable choice, which we can at least get a lower and upper bounds on in the human case (metabolic rates are pretty well studied, and even if we don’t have a good idea of how much energy is involved in completing cognitive tasks we can at least get bounds for running the entire system in that period of time) and close to a precise value in the ML case.