Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

921–930 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#921

Earlier quoted context omitted.

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

> They're packing lots of signal into fewer words There's a huge difference between the kind of prose you see in final output vs CoT windows. The final output is very much not what I'd call "packing lots of signal into fewer words" (aside perhaps from "Claude-isms" being easy enough to scan for if for some reason you actually wanted to scan for them, which other agents might want to for all I know); and if agents are…

I’ve lost track of the number of times I’ve told it to stop using terms like “evidence boundary” when writing specs. I still have no idea what that means.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#922
post #304
post #248

Earlier quoted context omitted.

Now that it's a solved benchmark, can we get the animated version?

I didn't want to shell out for Max again, so I piped the SVG created by Max back into Fable 5.1 at its default thinking level (of high): llm logs -cx | llm -m claude-fable-5.1 -s 'animate this' Here's the result, which cost $1.37: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... It's excellent!

Not sure if it's a rendering artefact, but the wheels are spinning backwards for me?

Also, 10 tooth sprocket with 22 tooth chainring? Not impossible, but so small!

Re: Claude Fable 5.1 and Claude Mythos 5.1

#923

Earlier quoted context omitted.

Qwen is all you need.

Not coincidentally, Claude is all Qwen needs.

NOOOOOOOOOOO, Chinees models are better than US models. They are cheap and open source. Which is the true blessing for humans. Otherwise the US would be the dictators for AI.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#924

Earlier quoted context omitted.

> as many of noted please rephrase?

"as many have noted", I suppose. I'm always baffled at how many people write "of" instead of "have", they don't even sound the same

Not sure where you're from but in my dialect (North American) it's more common than not to have _have_ realized as [əv] ("uhv") in contexts like _should have_, _could have_ (but not _I have a car_, where it has to be the full [hæv]). Only in deliberately enunciated speech do I feel like I'd expect [hæv] in the former kind of context. So it's an understandable mistake to make.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#925
post #248

Earlier quoted context omitted.

Now that it's a solved benchmark, can we get the animated version?

Not to be "that" person but it's not solved. The feet are reversed and isn't accurate bird anatomy. In real life, what people think of as bird's feet is actually their toes, and their "knee" is actually their tarsal (ankle bone), and their actual knee is almost hidden in their feathers.

The problem is it's kind of a question without a well defined answer. If it was just drawing a picture of a pelican, we can compare against real pelicans for accuracy. But what exactly is a pelican riding a bike supposed to look like? It's impossible, so liberties have to be taken somewhere.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#926
post #772

Earlier quoted context omitted.

Not sure if something like this is already on the table, but I would like to see Claude responses more in line with Simplified Technical English [1] by default. I find those writing styles a lot easier to read. This has been standardised as ASD-STE100 [2]. I've seen few people making SKILL.md files with that in mind, which works great, but having this by default without invoking the skill command would be better. 1.…

I must have missed something but can't you just prompt it to answer in your desired style? What am I missing here. Commenting because I am struggling with this too, claude code seems to be so verbose no matter how I prompt it.

Let's say Claude Code's system prompt is updated to recommend responding in Simplified Technical English. It'll work for the current generation of models. But Anthropic will train the next generation via RL on Claude Code traces as they do today. As long as they don't change their reward design, the next generation is going to be pulled towards Claudish again, because speaking Claudish gives higher rewards, so in the end the prompt doesn't really matter.

Ideally these things shouldn't work like this but this seems to be the sorry state of RL right now.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#927
post #46

Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…

Making clear the scale of distillation they’re combating.

It's only theft when people pay Anthropic for inference in order to improve their own datasets. It's not theft when Anthropic grabbed basically all ebooks and web content on the internet to build their own dataset, without paying anything to anyone

Re: Claude Fable 5.1 and Claude Mythos 5.1

#929

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

[flagged]

Re: Claude Fable 5.1 and Claude Mythos 5.1

#930

Earlier quoted context omitted.

Have you tried Grok 4.6, if you're focused on token budgets? In a league of it's own for tokens/intelligence.

SuperGrok quota is garbage for anything coding. I burn through my quota in a few hours with very mild use. SuperGrok Plus is slightly better but doesn’t last me more than a few days. Even Claude Max feels leagues more generous in usage… I haven’t tried SuperGrok Heavy because it’s too expensive

Yeah I actually prefer Grok 4.6 but the quotas are so low that I find myself using Codex 5.6 sol medium on their $100 plan.
Post reply on HN