Live data from Hacker News

I used to love Claude, but the latest models are slowly ruining it

androidauthority.com

51–60 of 63 posts

Re: I used to love Claude, but the latest models are slowly ruining it

#51
If the author is reading:

This is a probabilistic system. You’ve sampled from the distribution once per struggle then decided that was it, the model is just like this, Anthropic has it in for me...

But thankfully there’s a wide range of possibilities and (raw) LLMs don’t remember anything. So trying again may yield a different outcome. You have to try again.

Copy out the parts of the chat from before the trouble and bring them into a new chat, and start again from there.

Also. The contents of the chat are known to railroad the model. This means if a chat takes a turn toward suspicion, that’ll steer the model to more suspicion. If there is a refusal, there’ll be more likely to be more. So in your next attempt, rip out or don’t say whatever arose its suspicion.

And disable the memory feature. It wrecks anything scientific about these tools - for example: https://simonwillison.net/2025/May/21/chatgpt-new-memory/

If at first you don’t succeed, try again!

Re: I used to love Claude, but the latest models are slowly ruining it

#52
post #8

Claude has never been the best Chat agent. GPT and Gemini have the lead there. But Claude chat is still perfectly serviceable if you don’t want to pay for two.

> Claude has never been the best Hard disagree. It wasn’t that long ago that Gpt was clearly falling behind, and Gemini was like the “and also in the room”.

GPT was falling behind coding/work stuff vs claude, but I never felt it was a poorer generalist Chat agent.

I've used both quite a bit back/forth and this is just my personal opinion.

Re: I used to love Claude, but the latest models are slowly ruining it

#54
post #9

My own experience is that Opus 4.8 has an adversarial-teacher voice, unsolicited grading as if I submitted an essay for grading, declarations about the "real" issue, and constant "honest notes" self grading its own responses even before it answers. I can't stand its tone. We can't have a normal chat. While Fable reverts to Opus for simple questions like "What is digestion?"

The real issue with that tone (I'm contemplating cancelling my Anthropic subscription now) is that it's all too often the "confidently wrong teacher".

The tone is not the issue IMO: we're already at a point where we can have another, cheaper (as in: 20x cheaper for example), model query other models and have them reword the answers when having a "chat".

pi.dev can definitely control a Claude Code TUI window and follow prompts to Claude Code.

I'm 100% sure the arseholy tone of recent Claude Code models can be rewritten today, to not be arseholy, by literally prompting another model to unharsole Claude Code.

Now of course this doesn't solve the other half of the issue: in "teacher lecturing you while being confidently wrong", the "teacher lecturing you ..." is really not the most important issue.

Re: I used to love Claude, but the latest models are slowly ruining it

#55

You've got to understand, proprietary models operate as the following: - release: full precision, debrided, uncapped context - shortly after, hooked: quantized, governance department slammed, and a pseudo large context, attention reduced to start and end of thread. - down the road: 4bit quantized or worse with nerf incantation to make the next upcoming model feel amazing. Rince and repeat.

They have to be doing some variant of this in the life cycle because of how wildly the performance changes over time. I refuse to give Anthropic my own money because the subscription plans are essentially useless. It's rather nice with a company API account with no spending limit though. But really oAI models are more consistent over time in my experience.

Re: I used to love Claude, but the latest models are slowly ruining it

#56

Sol is really good

I'm tempted to try it out. I'm not keep to move but Fable rejected some work i was working on recently and frankly it's infuriating lol. I've been on Claude x20 for like 8 months now and now i'm tempted to switch out of spite.

It is surprisingly offensive.

The only friction for me is the general expensiveness of trying out top tier models, eg OpenAI's Fable equivalent (Sol?) to run for a trial period. I'd like to see a like-like comparison, eg buy x20 on OpenAI and see how much Sol i can use, how well it works, etc.

edit: Though surprisingly Claude seems to think Codex doesn't have hooks? That'll be tough, i use Claude hooks quite a bit.

Re: I used to love Claude, but the latest models are slowly ruining it

#57

This is a serious suggestion, not a joke: Have you tried being nice to the model? There are so many criticisms here that I just don't see myself. If the models have been trained on human responses, then it's plausible that they will prefer to become less helpful to requests which are blunt or even rude, because that's what humans do too.

I tend to anthropomorphize it as well, and keep reminding myself it's not human. It's not even a robot. Robots have routines, they follow what's burned in and don't ignore their config, outside of the Murderbot fantasies. But Claude (or Opus) ignore them routinely. I fall into asking it in the "could you please ..." manner but that feels wrong. I don't know what the "right" is though. Also when it or others reply "I am sorry" etc, well, these things are not an "I", but what are they?

Re: I used to love Claude, but the latest models are slowly ruining it

#58

Sol is really good

I'm tempted to try it out. I'm not keep to move but Fable rejected some work i was working on recently and frankly it's infuriating lol. I've been on Claude x20 for like 8 months now and now i'm tempted to switch out of spite. It is surprisingly offensive. The only friction for me is the general expensiveness of trying out top tier models, eg OpenAI's Fable equivalent (Sol?) to run for a trial period. I'd like to see…

It has them. It just lacks some of the features

Re: I used to love Claude, but the latest models are slowly ruining it

#60
post #9

My own experience is that Opus 4.8 has an adversarial-teacher voice, unsolicited grading as if I submitted an essay for grading, declarations about the "real" issue, and constant "honest notes" self grading its own responses even before it answers. I can't stand its tone. We can't have a normal chat. While Fable reverts to Opus for simple questions like "What is digestion?"

Yes, Opus 4.6 and Sonnet 4.6 are still Anthropic’s best chat models. They had more human evals and are still a joy to chat with. Opus 4.8 and Sonnet 5 by comparison have a cold enterprise feel. Opus 4.7 talks like I’d imagine a hitman would talk.

Sonnet 4.6 is better at writing good emails than Opus 4.8, Fable 5, and even Sonnet 5.

Post reply on HN