Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

911–920 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#912

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

It’s more likely that they have llms supervising llms in training and therefore the quality has dropped like a picture of a photograph.

If opus has high signal thinking it would be able to write a fsm but it’s been a month of me trying whereas Luna can do it in a few minutes.

I think it is similarly that they are using too much synthetic data.. meaning they are feeding the models the transcripts of users where many users have figured out to let agents just message each other.

Again picture of a photograph.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#914
post #405

I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets. I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited. I urge Anthropic to get be…

I literally only make it halfway through the week until my weekly usage runs out. This is using only Opus, no fable, and I'm on the max x20 plan. It's become ridiculous.

If you just say that you run out of tokens, it does not mean anything about the token quotas themselves being reasonable or not. That depends on how much you use it.

For instance, if you had 10s of agents running all the time, it is not that unexpected that you run out of tokens quickly.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#915

Earlier quoted context omitted.

Does it fix my favorite pet peeve, the overuse of the wrong meaning of "fail closed"? "Fail open" usually refers to a fuse that opens and kills power, meaning the system is inert and safe on failure. "Fail closed" is the opposite -- system has power and is live. Computer security people have appropriated the term but use it for the completely opposite meaning. When your work straddles electrical engineering and compu…

> fail closed I understand fail closed to mean, be secure when in failure. And fail open to be continue to operate during a failure. A door that fails closed would not let anyone in; one that fails open lets everyone in. But I can see how these are not the mutually exclusive definition the labels imply, especially if you apply the concept to entities that aren't doors or otherwise have explicit open/closed states. It…

Is there colloquial usage? It's common among computer programmers which is where Claude picked it up, but it's still an engineering term there.

It's comparable to "literally", which has picked up two opposite meanings, one of which appears to be obviously wrong. But because of the opposite meanings, using it the "correct" way is still wrong -- the only way to win is to not play, to stop using the terms "literally" and "fail-closed".

Re: Claude Fable 5.1 and Claude Mythos 5.1

#916

The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M). This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general. Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is h…

> Anthropic did not get much bite on Fable at its original pricing I stopped using Fable because it kept stopping itself due to safeguards.

Do you pay to get told no by Fable?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#917

Earlier quoted context omitted.

> People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths That people ARE making, surely not machines. Like Terence Tao or Knuts did using the tool to their advantage, for example it would have been impossible for me to prove the same thing Tao did with an LLM. Same reason I believe programmers won't go away

This isn’t true in math; see the proof of Crouzeix’s conjecture which was done by GPT 5.6 Sol in response to a prompt from a neurosurgery resident who had no deep math background, was learning that subject to better understand radiology, and thought it sounded like a cool theorem.

> Jin reported that the proof was obtained with the assistance of OpenAI's GPT-5.6 Sol model during an approximately sixteen-hour autonomous reasoning session in ChatGPT Work, after which he checked the resulting argument

You're wrong

Re: Claude Fable 5.1 and Claude Mythos 5.1

#919
post #405

I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets. I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited. I urge Anthropic to get be…

> I am becoming dependent on AI to make a living IMO, if you depend on AI to make a living, I'd invest in hardware for local inference, and learn on how to effectively make a living using AI inference you control, on hardware you control. Sure, economically speaking it's way cheaper to use one of these heavily subsidised services (for now), and their models are faster and more capable, but if your livelihood depends…

I somewhat agree, especially because things are still very dynamic. Who knows what access to meaningful amount of usage looks like in 1 year.

Unfortunately open models are still not as good as frontier ones and the hardware costs for a similar experience are very high.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#920
post #46

Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…

So basically, Anthropic can charge you for tokens you don't even see. "Trust me bro, you really did use $5,000 worth of tokens to generate that pelican". We need AI consumer rights, urgently.

I think that it would be sufficient to realize that no one is forced to use a specific AI model or model harness that has anti-consumer features built in.

People should just walk away if they see something like that. This is where the actual, practical, consumer rights start. Not with the regulation. But with the customers being determined to stand up for themselves. And not just fold.

There are plenty of good enough models that are open weight and whose use comes with almost no strings attached. And plenty of great harnesses such as various flavours of Pi.

Post reply on HN