> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.
You have a serious engineering problem if you're not able to find the source of a crash after years.
Claude Fable 5.1 and Claude Mythos 5.1
751–760 of 1001 posts
Re: Claude Fable 5.1 and Claude Mythos 5.1
#752Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#753Having lived through Covid, this doesn't sound so good to me.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#754Earlier quoted context omitted.
Do you think model trainers are pelicanmaxxing now?
There was a real probe into this, seems like no: https://dylancastillo.co/posts/pelicanmaxxing.html https://news.ycombinator.com/item?id=49010129
It'd be better to just have it draw a completely, linguistically, unrelated scene.
Like, a single tree in a meadow bending in the wind.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#755Earlier quoted context omitted.
A pelican is a bird, not a person. The knees bend the other way.
I get what you mean, but those are the ankles. The knees are higher up and bend the same way as a person's. https://en.wikipedia.org/wiki/Bird_feet_and_legs
A bit more direct
Re: Claude Fable 5.1 and Claude Mythos 5.1
#756I let it go a few hours on a not trivial but well-known problem, and it felt like it was just a little too plodding and just kind of mucked around a little too much and wasn't aggressive enough about getting stuff done. I asked it to wind it down and finish up and it took another hour and 15 to actually stop and commit without really getting much more done. Not very impressed here, if you can't tell. This new version…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#757Re: Claude Fable 5.1 and Claude Mythos 5.1
#758I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets. I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited. I urge Anthropic to get be…
I literally only make it halfway through the week until my weekly usage runs out. This is using only Opus, no fable, and I'm on the max x20 plan. It's become ridiculous.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#759Re: Claude Fable 5.1 and Claude Mythos 5.1
#760I'm confused about Anthropic's pricing. Can anyone explain why Sonner 5 is $2/MTok in and Sonnet 4.6 is still $3?
They originally released it at a "temporary discounted price", then made it permanent (probably due to competitive pressure). It's still way more expensive per task, due to tokenizer changes and general verbosity.
What would you suggest as an alternative if I'm happy with the harness?