Live data from Hacker News

Anthropic's best AI model struggles to attract users as cheaper tools thrive

ft.com

481–490 of 740 posts

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#481
post #179

Of course I echo the universal: Opus 5 sucks. But what also sucks is still using 4.8. Are you all seeing this? It's like the older models got dumber just before Fable and Opus 5 were coming out? I've heard the theory that it's because 4.8 is getting put on older hardware? And I imagine that just ratchets down reasoning time then possibly? Sometimes I'm just trying to sus out if I'm truly seeing things these days or g…

I'm not an expert by any means whatsoever but deploying models is not a straightforward task. There's a lot of levers to pull and I bet when models get "downgraded" to older hardware they do so WITHOUT the same stringent quality control of the output as they do when they release it. I don't think it's something deliberately malicious like planned obsolescence but it's more like startup culture of "just make it fit in…

I attribute intent. These are for-profit companies with more leveraged capex than any other industry in history. They have enormous pressure to optimise their limited compute. Especially Anthropic, which is far more limited than OpenAI. Of course they're pulling those levers in the background to achieve "good enough." If they can reduce memory consumption by 20% and their metrics show a 4% loss of intelligence, they might very well pull that lever. These decisions compound.

Your explanation is probably the most likely and largest contributor. Anthropic states that Opsu 5 is a "pinned snapshot". They claim the weights and model configuration are not silently updated, BUT the surrounding serving infrastructure can change, including the request router, safety classifiers, and sampling logic. Anthropic has stated that if behaviour unexpectedly changes on a stable model ID, an infrastructure update is the most likely cause.

Further, "High" isn't a fixed amount of compute. Anthropic describes effort as a "behavioural signal" and not a token budget, with the model deciding how much thinking to do. So their "High" might be "Low" now, and we would never know.

Finally, I strongly suspect some quantisation or KV-cache compression is happening. Anthropic doesn't clearly delineate whether this would fall under the pinned weights and configuration, or the infrastructure, which almost certainly guarantees it's the latter. Forgetting earlier information, poor retrieval of details, contradicting previous conclusions, hallucination, degraded instruction-following, and losing the thread during complicated tasks are all symptoms of quantisation and compression.

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#482
post #179

Of course I echo the universal: Opus 5 sucks. But what also sucks is still using 4.8. Are you all seeing this? It's like the older models got dumber just before Fable and Opus 5 were coming out? I've heard the theory that it's because 4.8 is getting put on older hardware? And I imagine that just ratchets down reasoning time then possibly? Sometimes I'm just trying to sus out if I'm truly seeing things these days or g…

RIGHT!? The data suggests this isn't true but every fibre of my being is convinced that Opus 5 High was EXCELLENT at launch and has been lobotomised since then. My benchmark is Sol High. I've been using both consistently and either Sol High suddenly became MUCH more capable - and the data does not support that - or Opus 5 became much dumber. It's so bad I can't even use it anymore.

IMO Opus 5 wasn't that great at launch, but 4.8 was definitely downgraded just prior.

Agree that Sol is a great model. I find that it's improved a bit but I mostly attribute this to its eagerness to use the harness' memory features. (I'm using it in Hermes Agent, FWIW)

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#483
post #295

Wow the sentiment here is so negative. I'm on the $200 plan (work pays) and I also have the $20 OpenAI plan (I pay) and keep a balance on OpenRouter. There is nothing as good as Fable, not even close. I recently had it run a 18 hour autonomous rebuild of a project (moving from Spark to Pandas for performance/data size trade off issues). It orchestrated Opus sub-agents flawlessly for 18 hours. It even did a great job…

[deleted]

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#484
post #295

Wow the sentiment here is so negative. I'm on the $200 plan (work pays) and I also have the $20 OpenAI plan (I pay) and keep a balance on OpenRouter. There is nothing as good as Fable, not even close. I recently had it run a 18 hour autonomous rebuild of a project (moving from Spark to Pandas for performance/data size trade off issues). It orchestrated Opus sub-agents flawlessly for 18 hours. It even did a great job…

Like others have suggested, you should give more time to GPT models. I sometimes launch Fable with elaborate review personas, it might take 30 minutes or an hour, exceed limits, to produce a review of a PR. Then I ask the same thing GPT without any elaborate 'come up with personas, review the reviews, do rebuttals, etc', and it can find problems that hours of Fable couldn't.

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#485

Non-LLM user here. Why? Apart from the ecological issues, I'm very uncomfortable giving my organisation's crown jewels to {random_internet__corp}. Look at the lengths they go to for training data - 10M for Spirit's call logs? Destroying millions of obscure books to scan them? They make meth-heads look scrupulous. Your prompts, especially if they contain your entire codebase, are _way_ more data-rich than an airline p…

creating mvp

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#486
post #113

Where Anthropic f'ed up was treating their monetization the way they treat model training. Turns out that success in experimentation is not transferrable. They have tried to find the highest that the market pays for sota models; however, on the consumer side, this is just too confusing and unsettling: "You can only use Fable for a week as a part of your plan" "Be ready! You have to start paying per token!" "Nevermind…

> That forces people to look beyond the walled garden. Every time I get "you used your quota, come back in 3 hours, or 2 days" -> that is experimentation time with their competition, leading to changed service plans. When they said "claude -p" will be billed at API pricing even for plan users I moved my harness off claude. After I integrated codex, then it was never going to be a full claude project again. What busin…

> What business encourages users to try their competition and adapt their usage to the competing products?

They are high on their own supply. The people running these companies are delusional imbeciles who have been placed in charge of billions of dollars.

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#487

Earlier quoted context omitted.

As soon as I started using Fable I was like, okay, this is probably as good a model as I will need for software engineering going forward. I still feel that way. I don’t need a better model, I need a faster Fable. The thing I miss most about programming is flow, and the constant bouncing between terminal tabs sucks. I’d love to do one thing at a time, with Fable, quickly.

remember 4 year ago we use to : have stack overflow open, documentation, obscure forums plus other tabs. An ide open with 20 tabs open each file a component, a class or an interface We also use to hold entire codebases in our brain.

Yeah StackOverflow which was either telling you to use google or it was so specific, that no one wanted/could respond.

Even a year ago when i was trying to do a hugo template manually with the help of the documentatin /tutorial, it was shit. The LLM at that time, was better helping me than the documentation.

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#488
post #113

Where Anthropic f'ed up was treating their monetization the way they treat model training. Turns out that success in experimentation is not transferrable. They have tried to find the highest that the market pays for sota models; however, on the consumer side, this is just too confusing and unsettling: "You can only use Fable for a week as a part of your plan" "Be ready! You have to start paying per token!" "Nevermind…

It also distorts the testing as it encourages non-typical behavior.

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#489
The reading seems a little presumptuous to me. We do not know what Anthropic expected. This is mostly a matter of performance/cost and they clearly focused on building the Rolls Royce. How much adoption do you expect, when you build the Ferrari? And how many Rolls Royce can you even deliver?

Saying "people are not buying that many Rolls Royce, instead they buy a lot of normal cars" is kind of duh. Anthropic is clearly still operating at inference capacity.

Re: Anthropic's best AI model struggles to attract users as cheaper tools thrive

#490
post #408

Non-LLM user here. Why? Apart from the ecological issues, I'm very uncomfortable giving my organisation's crown jewels to {random_internet__corp}. Look at the lengths they go to for training data - 10M for Spirit's call logs? Destroying millions of obscure books to scan them? They make meth-heads look scrupulous. Your prompts, especially if they contain your entire codebase, are _way_ more data-rich than an airline p…

I’m not certain of that - I feel like the enterprise zero retention policy will be tight?

[dead]
Post reply on HN