Earlier quoted context omitted.
Fable is twice the price.
If Fable gets correct answer quicker, then you might pay less than doing back and forth with Opus, plus you lose more of your own time. I see no reason for using less able models in my workflows. There is this saying, penny wise and pound foolish
Claude Opus 5
191–200 of 1001 posts
Re: Claude Opus 5
#192Annoyingly, this is a concrete argument that open source software may be easier to attack.
Re: Claude Opus 5
#193Re: Claude Opus 5
#194That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?
Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores. I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.
Re: Claude Opus 5
#195Re: Claude Opus 5
#196Where does this leave Fable? I am confused.
Re: Claude Opus 5
#197Earlier quoted context omitted.
Funny that a company selling an AI software developer can't use it to fix their infra. Fixing those issues still requires humans.
Let's be honest - they're also still hiring software devs. AI still requires skilled humans in the loop and that's not going away.
Imagine you are a company that sells concrete. You have a web dev contractor you use to build and maintain your website. It has tools on it to get delivery quotes and a few internal tools to track orders.
Except now you can just have your sales team also maintain the website with a $20/month Claude subscription.
Re: Claude Opus 5
#198Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
why is that? its now being benchmaxxed too
Re: Claude Opus 5
#199Is Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?
I suspect they make a big model first. In this case it's Fable. Then they run the shrinker steps to make Sonnet and Opus. Sonnet is smaller, takes less time to make, so it got released first. Opus needed few more weeks to cook. With this iteration they had a delay because when the Mythos was ready they had some sort of "Oh shit" moment and spent half a year adding safety guards to it. Then slowly rolled it out, but g…
Re: Claude Opus 5
#200Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
seeing a jump this big is not a great sign for the continuing value of a benchmark