Live data from Hacker News

Anthropic's best AI model struggles to attract users as cheaper tools thrive

ft.com

41–50 of 740 posts

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#41
As a small background, I have a local server and I've been trying out different models with different inference engines, quants, configurations etc... I'm also using Opus and Sol at work consistently. I've used AI since the first wave, first as a toy, then as a highly specific tool, last 6+ months as the primary LoC generator.

This is the first time I've felt, and I use the word *felt* since I don't have a suite of benchmarks or any sort of material approach towards comparing models, that Opus has declined in quality compared to before. Primarily I think its powers of deduction and understanding, even on xhigh, have become much worse. Before, being vague and providing a simple prompt would be enough, it could deduce and expand the details it needed, plus ask you clarifying questions, now this is no longer the case. A concrete, personal example, for a personal project, I've asked it to setup ssl over local IP. I didn't go into too much detail in the prompt as there are many approaches it could take and I didn't care too much to choose. It did horrible. The first thing it did was say the best lightweight approach is to add a reverse proxy. I'm like ok, makes sense. Then after asking it to proceed, it went and added a bunch of config to my golang service and didn't even setup a reverse proxy even when it said that is the way to go. It even said it didn't set it up lol. Then after I told it to do so it failed building the config in a way it was asked of it (support LAN IP and tailscale IP). Etc etc...

When Fable came out it was huge, the benchmarks told the story, and the story mostly matched the experience. It felt, again, intentionally saying felt, like it was miles ahead. Now benchmarks say that there are many models that are close, but in actual use Fable still *feels* much better. I think benchmaxxing the new open weights models is ruining the value of benchmarks, if they ever had any. When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.

The price for Fable is definitely too much for any personal use now that it's no longer included in the subscription, and GLM 5.2, Deepseek Flash and Qwen 3.8 served locally or via cloud provide a lot, requiring a bit more babysitting though. Considering the price of Fable, my 5k USD Epyc server would pay itself off in less than a year if I used Fable or Opus in the same manner so at least for me the decision seems easy. And considering the point I'm poorly trying to make, that Opus doesn't feel like frontier anymore, this is probably the last month of my Claude subscription.

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#42
post #36

Anthropic’s issue is churn because of the peak verbosity vomit coming out of Opus 5/Mythos/Fable. What the hell did they train it on. The sane model is still Opus 4.6.

Perhaps there is a sophisticated subtle poisoning attack that makes models behave like that?

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#43
post #27

They've put themselves in a corner. Fable was too good and they gave it away with the $20 plan. It had to be a big step from Opus 4.8 to show progress, and Opus 4.8 is GREAT at coding in many different domains. But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _…

Fable is still on the $100 plan for me. (Maybe this is A/B testing or something.)

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#44
post #35

Earlier quoted context omitted.

[flagged]

Yes that is exactly what's happening. Claude keeps track of memory if you've turned that on so all conversations are somehow tracked over time Claude randomly will mention that I'm a developer while I'm asking unrelated questions and say oh because you're a developer you might like this or because of my background and infrastructure you might find this interesting and I always get creeped out by it. They have the con…

[flagged]

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#45
post #7

Fable is not a tool for the average user. It’s a professional tool for highly complex work. I would compare it to a extremely high end $15k PC, or an expensive pro-grade video camera, or a freight train, or a … I would say at least 95% of the global population will not encounter a situation once in their life where it would be actually useful/warranted.

I wish Fable were as good as you make it sound. A plan created by Fable is good, but in my case, it always contain issues caught only when it's reviewed again (whether by itself, Opus, Sol etc.). That's (almost) not different from plans created by Sol, GLM 5.3 etc. The one thing where it's genuinely better is the front-end, but then again it's far from perfect, it just needs less iterations.

>> but in my case, it always contain issues caught only when it's reviewed again

Yes but those issues will be much less severe with Fable-written plans than those written by lesser models. I know this because my workflows at both my regular job and my startup involve multi-step agent reviews via codified adversarial review skills. Fable as a reviewer will frequently find blocker-level issues with plans written by GPT 5.6 Sol, and sometimes with Opus 5. The opposite almost never happens. In fact I cannot remember the last time it happened.

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#46
post #12

Earlier quoted context omitted.

Just cancelled my Claude subscription and used Ox Alpha Free through OpenCode (so that's one). In my view, from the testing I did since Thursday, it's better than Fable. I had just finished a rather large task that Fable completed, including a /review and an /ultrareview. Ox Alpha found bugs that Fable and Opus missed, and it continued to build things like a pro. It does have issues with availability - but it's on a…

Intriguing. Who is it? What is it? any ideas? https://openrouter.ai/stealth/ox-alpha

Rumored to be mimo

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#47
Karma backed over their dogma.

I switched from Opus 4.7/4.8 to test Kimi K3 a few weeks back and the test hasn't finished; it's my daily driver now.

Given their general behavior and preference toward social engineering to scare the shit out of normal people...this seems fitting.

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#48
post #27

They've put themselves in a corner. Fable was too good and they gave it away with the $20 plan. It had to be a big step from Opus 4.8 to show progress, and Opus 4.8 is GREAT at coding in many different domains. But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _…

[deleted]

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#49
post #27

They've put themselves in a corner. Fable was too good and they gave it away with the $20 plan. It had to be a big step from Opus 4.8 to show progress, and Opus 4.8 is GREAT at coding in many different domains. But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _…

[deleted]

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#50
post #7

Fable is not a tool for the average user. It’s a professional tool for highly complex work. I would compare it to a extremely high end $15k PC, or an expensive pro-grade video camera, or a freight train, or a … I would say at least 95% of the global population will not encounter a situation once in their life where it would be actually useful/warranted.

You have a $3 trillion bubble riding on this not being true.

The bubble is _not_ on models becoming more intelligent and solving arc-agi-999.

They are already good enough at what they mechanically are.

You have to use the right harness, right verifiers (automatic where possible, human where not), etc much much more specific than a generic one like claude code or codex, and it will also be able to work within constraints and be the "proposer" of an imaginary optimisation problem and an excellent one at that. But you have to frame the task at hand in that manner or maybe even reorganise the task you do itself so it is more amenable to being framed that way. If you use it this way, it is _already_ massively economically useful. But it will take many years for it to actually be usable in that way, since you need DC capacity to come up first which is few years away and also well, massive organisations that have to integrate these will usually take many years to do so.

It is also useful albeit less so in cases like general SWE, where you still need a human in a loop for non-verifiable requirements, and also in other general usecases where information retrieval is too intractable and you need to carefully use LLMs as a component of the overall system.

I am not saying Fable or whatever the biggest models are are useless - they will certainly be useful for tasks at the frontier of the day - which is today complex exploits and open math problems, and well, tomorrow it could be something in biotech. But this is not what the entire bet is on at all - just automating day to day drudge at the tens of thousands of massive companies and governments we all know and love is more than enough. With the right training data (which _also_ is a bottleneck and takes time) you could even automate certain processes entirely. Sure, if we get a crazy medical innovation and end up saving trillions in healthcare great, but that's just a bonus.

None of this is to say that I think there is zero sketchy financial engineering going on

Post reply on HN