Live data from Hacker News

Claude 4

anthropic.com

351–360 of 1001 posts

Re: Claude 4

#351

I've been using Claude Opus 4 the past couple of hours. I absolutely HATE the new personality it's got. Like ChatGPT at its worst. Awful. Completely over the top "this is brilliant" or "this completely destroys the argument!" or "this is catastrophically bad for them". I hope they fix this very quickly.

When I find some stupidity that 3.7 has committed and it says “Great catch! You’re absolutely right!” I just want to reach into cyberspace and slap it. It’s like a Douglas Adams character.

Re: Claude 4

#352
post #206

Earlier quoted context omitted.

With web search being available in all major user-facing LLM products now (and I believe in some APIs as well, sometimes unintentionally), I feel like the exact month of cutoff is becoming less and less relevant, at least in my personal experience. The models I'm regularly using are usually smart enough to figure out that they should be pulling in new information for a given topic.

It still matters for software packages. Particularly python packages that have to do with programming with AI! They are evolving quickly, with deprecation and updated documentation. Having to correct for this in system prompts is a pain. It would be great if the models were updating portions of their content more recently than others. For the tailwind example in parent-sibling comment, should absolutely be as up to d…

> the history of the US civil war can probably be updated less frequently.

It's already missed out on two issues of Civil War History: https://muse.jhu.edu/journal/42

Contrary to the prevailing belief in tech circles, there's a lot in history/social science that we don't know and are still figuring out. It's not IEEE Transactions on Pattern Analysis and Machine Intelligence (four issues since March), but it's not nothing.

Re: Claude 4

#353
post #51

I've found myself having brand loyalty to Claude. I don't really trust any of the other models with coding, the only one I even let close to my work is Claude. And this is after trying most of them. Looking forward to trying 4.

I'm slutty. I tend to use all four at once: Claude, Grok, Gemini and OpenAI.

They keep leap-frogging each other. My preference has been the output from Gemini these last few weeks. Going to check out Claude now.

Re: Claude 4

#354
post #311

Earlier quoted context omitted.

Guess we have to wait till DeepSeek mops the floor with everyone again.

DeepSeek never mopped the floor with anyone... DeepSeek was remarkable because it is claimed that they spent a lot less training it, and without Nvidia GPUs, and because they had the best open weight model for a while. The only area they mopped the floor in was open source models, which had been stagnating for a while. But qwen3 mopped the floor with DeepSeek R1.

I think qwen3:R1 is apples:oranges, if you mean the 32B models. R1 has 20x the parameters and likely roughly as much knowledge about the world. One is a really good general model, while you can run the other one on commodity hardware. Subjectively, R1 is way better at coding, and Qwen3 is really good only at benchmarks - take a look at aider‘s leaderboard, it’s not even close: https://aider.chat/docs/leaderboards/

R2 could turn out really really good, but we‘ll see.

Re: Claude 4

#355

Earlier quoted context omitted.

What's with all the models exhibiting sycophancy at the same time? Recently ChatGPT, Gemini 2.5 Pro latest seems more sycophantic, now Claude. Is it deliberate, or a side effect?

IMO, I always read that as a psychological trick to get people more comfortable with it and encourage usage. Who doesn't like a friend who's always encouraging, supportive, and accepting of their ideas?

I want friends who poke holes in my ideas and give me better ones so I don't waste what limited time I have on planet earth.

Re: Claude 4

#356
context window of both opus and sonnet 4 are still the same 200kt as with sonnet-3.7, underwhelming compared to both latest gimini and gpt-4.1 that are clocking at 1mt. For coding tasks context window size does matter.

Re: Claude 4

#357

Earlier quoted context omitted.

What's with all the models exhibiting sycophancy at the same time? Recently ChatGPT, Gemini 2.5 Pro latest seems more sycophantic, now Claude. Is it deliberate, or a side effect?

I think OpenAI said it had something to do with over-indexing on user feedback (upvote / downvote on model responses). The users like to be glazed.

If there's one thing I know about many people (with all the caveats of a broad universal stereotype of course), they do love having egos stroked and smoke blown up their ass. Give a decent salesperson a pack of cigarettes and a short length of hose and they can sell ice to an Inuit.

I wouldn't be surprised at all if the sycophancy is due to A/B testing and incorporating user responses into model behavior. Hell, for a while there ChatGPT was openly doing it, routinely asking us to rate "which answer is better" (Note: I'm not saying this is a bad thing, just speculating on potential unintended consequences)

Re: Claude 4

#358

Earlier quoted context omitted.

(Caveat on this comment: I'm the COO of OpenRouter. I'm not here to plug my employer; just ran across this and think this suggestion is helpful) Feel free to give OpenRouter a try; part of the value prop is that you purchase credits and they are fungible across whatever models & providers you want. We just got Sonnet 4 live. We have a chatroom on the website, that simply uses the API under the covers (and deducts cre…

Just wanted to provide some (hopefully) helpful feedback from a potential customer that likely would have been, but bounced away due to ambiguity around pricing. It's too hard to find out what markup y'all charge on top of the APIs. I understand it varies based on the model, but this page (which is what clicking on the "Pricing" link from the website takes you to) https://openrouter.ai/models is way too complicated.…

From https://openrouter.ai/docs/faq#how-do-i-get-billed-for-my-us...

> We pass through the pricing of the underlying providers; there is no markup on inference pricing (however we do charge a fee when purchasing credits).

Re: Claude 4

#359

Earlier quoted context omitted.

What's with all the models exhibiting sycophancy at the same time? Recently ChatGPT, Gemini 2.5 Pro latest seems more sycophantic, now Claude. Is it deliberate, or a side effect?

It’s starting to go mainstream. Which means more general population is given feedback on outputs. So my guess is people are less likely to downvote things they disagree with when the LLM is really emphatic or if the LLM is sycophantic (towards user) in its response.

The unwashed hordes of normies are invading AI, just like they invaded the internet 20 years ago. Will GPT-4o be Geocities, and GPT-6, Facebook?

Re: Claude 4

#360

I've been using Claude Opus 4 the past couple of hours. I absolutely HATE the new personality it's got. Like ChatGPT at its worst. Awful. Completely over the top "this is brilliant" or "this completely destroys the argument!" or "this is catastrophically bad for them". I hope they fix this very quickly.

It's interesting that we are in a world state in which "HATE the new personality it's got" is applied to AIs. We're living in the future ya'll :)
Post reply on HN