Frontier LLMs drop from 83% to 43% once reasoning has to chain across domains
1–2 of 2 posts
Re: Frontier LLMs drop from 83% to 43% once reasoning has to chain across domains
#2Is that the model, all models, the agent?
Does something like this make a difference?
"Schema Harness Achieves ~99% on Arc‑AGI‑3 Public" (2026-07) https://news.ycombinator.com/item?id=48935905