A lot of skepticism here, but these are astonishing results! People should realize we’re reaching the point where LLMs are surpassing humans in any task limited in scope enough to be a “benchmark”. And as anyone who’s spent time using Claude 3.5 Sonnet / GPT-4o can attest, these things really are useful and smart! (And, if these results hold up, O1 is much, much smarter.) This is a nerve-wracking time to be a knowled…
> People should realize we’re reaching the point where LLMs are surpassing humans in any task limited in scope enough to be a “benchmark”. Can you explain what this statement means? It sounds like you're saying LLMs are now smart enough to be able to jump through arbitrary hoops but are not able to do so when taken outside of that comfort zone. If my reading is correct then it sounds like skepticism is still warrante…
We are seeing steady improvement on long-run tasks (SWE-Bench being one example) and much more improvement on shorter, more well-defined tasks. The latter capabilities aren’t “hype” or just for show, there really is productive work like that to be done in the world! It’s just not everything, yet.