Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

751–760 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#751
post #144

Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…

The biggest issue i found was the gap between the tire and rim. Otherwise, the max version is the best so far!

Re: Claude Fable 5.1 and Claude Mythos 5.1

#753

Earlier quoted context omitted.

Do you think model trainers are pelicanmaxxing now?

There was a real probe into this, seems like no: https://dylancastillo.co/posts/pelicanmaxxing.html https://news.ycombinator.com/item?id=49010129

I don't really see how replacing the vehicle and the animal is a good test.

It'd be better to just have it draw a completely, linguistically, unrelated scene.

Like, a single tree in a meadow bending in the wind.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#754
post #681

Earlier quoted context omitted.

A pelican is a bird, not a person. The knees bend the other way.

I get what you mean, but those are the ankles. The knees are higher up and bend the same way as a person's. https://en.wikipedia.org/wiki/Bird_feet_and_legs

https://en.wikipedia.org/wiki/File:Bird_leg_and_pelvic_girdl...

A bit more direct

Re: Claude Fable 5.1 and Claude Mythos 5.1

#755

I let it go a few hours on a not trivial but well-known problem, and it felt like it was just a little too plodding and just kind of mucked around a little too much and wasn't aggressive enough about getting stuff done. I asked it to wind it down and finish up and it took another hour and 15 to actually stop and commit without really getting much more done. Not very impressed here, if you can't tell. This new version…

Compaction is a setting that you control in Claude Code.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#757
post #405

I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets. I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited. I urge Anthropic to get be…

I literally only make it halfway through the week until my weekly usage runs out. This is using only Opus, no fable, and I'm on the max x20 plan. It's become ridiculous.

just buy two 20x, no?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#759

I'm confused about Anthropic's pricing. Can anyone explain why Sonner 5 is $2/MTok in and Sonnet 4.6 is still $3?

They originally released it at a "temporary discounted price", then made it permanent (probably due to competitive pressure). It's still way more expensive per task, due to tokenizer changes and general verbosity.

Is it? I just did a test switch over. For my personal needs I set up a box with OpenClaw back in March, which feels like a million years ago, that's been running Sonnet 4.6 since then. With all the caching it seems like my actual cost has come out around $1 per MTok on that. I just updated my whole setup today to try Sonnet 5... so far it looks like it's using fewer tokens for similar tasks, but it's only been half a day. I'm not super interested in changing harnesses, I realize this might not be the cheapest way but I've sorta come to enjoy OpenClaw... it's relatively effortless and responsive, and brief, given full control of a machine. And it does seem to incur some significant savings with the way it manages to keep things cached.

What would you suggest as an alternative if I'm happy with the harness?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#760

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

By injecting that weird prompt and not by proper post training? anthropic is truly a joke.
Post reply on HN