Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

961–970 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#961
post #405

I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets. I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited. I urge Anthropic to get be…

> I am becoming dependent on AI to make a living IMO, if you depend on AI to make a living, I'd invest in hardware for local inference, and learn on how to effectively make a living using AI inference you control, on hardware you control. Sure, economically speaking it's way cheaper to use one of these heavily subsidised services (for now), and their models are faster and more capable, but if your livelihood depends…

This is pretty terrible advice when there are dozens of AI inference providers out there serving great models with significantly more cost effectiveness than you'd get from buying your own hardware.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#963

Earlier quoted context omitted.

Have you tried Grok 4.6, if you're focused on token budgets? In a league of it's own for tokens/intelligence.

Not for enterprise. Can't trust the company behind it with my data.

very ironic if you think you can trust open ai over xai with you data.

if you don't trust any, then at least that's a coherent position

Re: Claude Fable 5.1 and Claude Mythos 5.1

#964
post #867
post #243

I've recently been running these agent sessions on more and more long running tasks because these latest models can do a REALLY good job on big chunks of work, and i've been watching them way less. It's starting to occur to me the importance of alignment is a today problem, it's not a tomorrow problem. In the past I watched and saw everything the model did, not a lot got past me. Today it does A TON of work while i'm…

what do you do while agents work?

Run more agents!

Re: Claude Fable 5.1 and Claude Mythos 5.1

#965

The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M). This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general. Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is h…

I don't think it's a stall, two ways I would believe there is a stall:

- Does the epoch capability index progress show signs of plateauing? I consider this a good aggregate measure of diverse benchmarks into a single capability index. If we see things slowing down here thats a pretty direct and convincing piece of evidence for a stall. - Do we see any signs that scaling laws are beginning to fail? That would be by far the most alarming to me, since I would interpret that to mean that the entire premise of this unprecedented capital allocation tsunami is broken.

Neither of these are true (for now). Progress is marching the same as it has for 4+ years now. It's still the same time to get a generation leap (I think like ~16-18 mo? Epoch has it) like GPT4->5. My theory is that people interpret plateauing because the releases are far more frequent now than they were in the past.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#966
I'm late to the thread, but my experience with Claude Fable 5.1 has been absolutely horrendous.

Things it does constantly that Fable 5 barely ever did:

- Act without my permission. All. The. Time. "Oh I just finished this thing we were discussing, let me push it without ever having been told to do so."

- Immediately jump to action instead of addressing me first. If I say "I wanted to write tests for this and run them" it immediately starts writing tests instead of digging into what "this" is better -- literally does not give me any feedback and starts spitting out code. Naturally it creates the wrong tests

- Despite claims that it does not write like "stereotypical Claude" anymore, in my experiments it is far worse than before. Replies are longer, more filled with fluff, and still flooded with garbage language. Hard to parse.

- It loves to answer my set of two direct Yes/No questions with 5 paragraphs where it only answers one of them and answers 4 other questions I didn't ask. Notice how it misses one of the questions.

- It. Is. Cocky. Absurdly full of itself and arrogant. Just the whole way it presents and answers passes this energy of "No, but really, you're wrong and I'm right". It often is not right. What annoys me is not that it's wrong more often than before (which it may be), it's that it doesn't own up to it as before. Insulting if it were a human.

- Replies and addresses me directly in its thinking traces, and then assumes I've read it. I ask a question, it answers it in the thinking traces and does not relay it back to me at all. This is the only one that Fable 5 also did, but 5.1 is doing it an order of magnitude more often.

- It's too early to really tell, because I may just be working on particularly harder problems today, but it seems to get things wrong more often. I've had to bump it from high to xhigh to compensate.

My guess is I must be having a bad day or something. Although this is happening on multiple projects run from multiple machines (fully isolated, except for the account, which is the same) all in the same way.

Will probably downgrade to 5 while I can.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#967
post #468

Earlier quoted context omitted.

I'm a heavy user and fable is great the #1 reason I stopped using it was the horrible safegaurd filter. I found sol close enough in capability and have only been blocked when my request was an obvious offensive cyber work. Fable blocked me on almost everything. Optimizing a OS build? -> block Securing a container -> block 60% is nowhere near enough for that safegaurd system. This just means I am going to be blocked h…

I never hit Anthropic's safety filter when I'm doing something illegal, only when I'm not.

Mere mention of "reverse engineering" gets me kicked back to Opus.

Where I reside, reverse engineering for interoperability is generally legal, and interoperability (e.g. getting a USB HID and USB MIDI devices or DOS programs to work in Linux/Android) is essentially what I'm interested in.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#968

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

It's totally valid if the models want to pack words tight during their thinking process, as long as the final conclusion (which is the interface to user) is written in HUMAN LANGUAGE, then I don't care whether the model thinks in alien language

Re: Claude Fable 5.1 and Claude Mythos 5.1

#969

I am absolutely thrilled that they reset weekly limits. I have been experimenting with highly autonomous work (5+ hours continuous) and fable seems excellent at this, especially when using subagents. I ran out of Fable capacity and was bummed out that my experiment would take longer to complete. Now I'm super happy I get to continue it

No other model have been able to complete your highly autonomous work? None? Really? Sounds a bit dystopian to be thrilled about a weekly reset so you can continue to work.

My experiment is examining the autonomy of Fable specifically in an auto research context. I don't believe I said in my message that no other model would have been able to complete my highly autonomous work. So it feels like my view has been misrepresented or misunderstood. This message makes it harder for me to share the things that excite me online and makes it more daunting to share my findings when this project completes, especially any comparative work. For an analogy, I feel like I said that I like Southern Butter Pecan Ice Cream and am being met with a response of the form "Sounds a bit sad that you have to wait for a weekly restock to enjoy any ice cream." I made a goal for myself to be more open with my feelings in life and share more of what I'm working on and not be so rejection-sensitive. I understand that even if I'm just sharing the positivity I feel, it can come across differently. I guess this is just the cost of communication in a lossy language.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#970

Instead of a new model that's going to have unreasonably shallow usage limits, I wish they would: 1) address the claude 20x plan usage being only 6-7x the ceiling of the claude pro plan 2) either fix opus 5, make it completely free, or delete it entirely

I downgraded from the 20x today after learning that 20x only applies to 5 hour usage. I have barely used Claude/Claude Code in the last month and am considering downgrading further, even after this update.

Wait the 20x doesnt multiply the weekly?
Post reply on HN