Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

821–830 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#821
post #772

Earlier quoted context omitted.

Not sure if something like this is already on the table, but I would like to see Claude responses more in line with Simplified Technical English [1] by default. I find those writing styles a lot easier to read. This has been standardised as ASD-STE100 [2]. I've seen few people making SKILL.md files with that in mind, which works great, but having this by default without invoking the skill command would be better. 1.…

If someone has a way to tame GLM into doing this then please… As a workhorse, GLM is so good, but goodness, its prose, wherever needed, makes me feel like going to a park and kicking all the benches there endlessly. And it doesn't change!

I have instructions to combine style of Brooks, Steven King and Economist Style Guide, while avoiding anything that may put wrong emphasis („genuinely“, „load-bearing“) does not contribute anything meaningful („X rather than Y“). The output is quite good* so far even on Opus 4.7-5, when it concerns product requirements. The key is to identify the pattern and focus on the meaning of the undesirable language.

* chemistry + regulations + infosec domains

Re: Claude Fable 5.1 and Claude Mythos 5.1

#822
post #405

I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets. I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited. I urge Anthropic to get be…

Try gpt 5.6 Luna max

overthinks, been slow lately through the official api (slower than glm 5.3 somehow), and tries to run every conceivable e2e test once it does literally anything.

like yesterday it ran for like an hour to build a fairly basic frontend...

i like luna and sol but it feels bad lately

Re: Claude Fable 5.1 and Claude Mythos 5.1

#825
post #46

Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…

So basically, Anthropic can charge you for tokens you don't even see. "Trust me bro, you really did use $5,000 worth of tokens to generate that pelican". We need AI consumer rights, urgently.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#826

I let it go a few hours on a not trivial but well-known problem, and it felt like it was just a little too plodding and just kind of mucked around a little too much and wasn't aggressive enough about getting stuff done. I asked it to wind it down and finish up and it took another hour and 15 to actually stop and commit without really getting much more done. Not very impressed here, if you can't tell. This new version…

The incentives are not aligned. Anthropic makes money when you burn tokens. They also hide the tokens you burn from you, so you can't even validate if you actually burned them, you just have to believe them.

This is not a lasting business model, nor one I'm interested in using.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#827

> This required us to add a watermark—a numerical way of determining the likelihood that Claude was involved in writing a piece of text—to the outputs of models released after August 2, 2026. As we recently explained, this watermark is invisible to anyone who does not have the detection API. It has no practical impact on the quality or content of Claude’s outputs and contains no information about the user, their orga…

It doesn't.

The more realistic claim it does not affect output quality.

Anthropic and fanboys defend that choosing "overcast" over "cloudy" does not affect a text quality.

Clearly, money and ambition clouds their judgment.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#828

To be honest, these frontier model releases have become boring for me. Opus 4.8 was already good enough for most of my use cases. I don't have any projects right now that I would use Fable for instead of Opus. So when I see announcements like this I just think "that's cool I guess" and then go back to using weaker/cheaper models. What's far more exciting right now is models like DeepSeek V4 Flash and GLM 5.3 Flash. T…

Not to mention you can actually see the tokens you pay for, and you can finetune it or change its system prompt.

I think compute providers are the big business. Run any model you like, adjust the weights however you like, adjust the prompts and behavior however you like, but we'll provide all of the hardware infrastructure for you at a renting fee.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#829
post #587

Earlier quoted context omitted.

I'm with you, for what I usually do most models are already more than enough. What I'm really keen on is better auto-reasoning so I don't have to constantly have the constant inner debate on which reasoning effort to pick for each task. I seriously hate the none-low-medium-high-xhigh-max-ultra etc that we have now, with companies frequently recommending different ones on each new model release, etc. It's apparently c…

Adaptive reasoning is known to be an extremely hard problem to solve, though. It requires you to predict whether a certain LLM, with a certain effort level, with a certain prompt, will give you the right answer.

This feels like exactly the kind of problem domain that belongs in (and can be solved by) RL?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#830
post #452

All the benchmarks in the world don't matter if the model just straight up refuses to do mundane things. Claude has too much of an attitude.

I'm a kernel engineer. Fable 5 refused all my requests, falling back to Opus 4.8. My wife is a chemist. Her experience wasn't much better.

GLM 5.3 is supposedly incredible for kernel engineering. Have you tried it?
Post reply on HN