Live data from Hacker News

Claude 4

anthropic.com

341–350 of 1001 posts

Re: Claude 4

#341

Is this really worthy of a claude 4 label? Was there a new pre-training run? Cause this feels like 3.8... only swe went up significantly, and that as we all understand by now is done by cramming on specific post training data and doesn't generalize to intelligence. The agentic tooluse didn't improve and this says to me that it's not really smarter.

I was about to comment on a past remark from Anthropic that the whole reason for the convoluted naming scheme was because they wanted to wait until they had a model worth of the "Claude 4" title.

But because of all the incremental improvements since then, the irony is that this merely feels like an incremental improvement. It obviously is a huge leap when you consider that the best Claude 3 ever got on SWE-verified was just under 20% (combined with SWE-agent), but compared to Claude 3.7 it doesn't feel like that big of a deal, at least when it comes to SWE-bench results.

Is it worthy? Sure, why not, compared to the original Claude 3 at any rate, but this habit of incremental improvement means that a major new release feels kind of ordinary.

Re: Claude 4

#342

An important note not mentioned in this announcement is that Claude 4's training cutoff date is March 2025, which is the latest of any recent model. (Gemini 2.5 has a cutoff of January 2025) https://docs.anthropic.com/en/docs/about-claude/models/overv...

Should we not necessarily assume that it would have some FastHTML training with that March 2025 cutoff date? I'd hope so but I guess it's more likely that it still hasn't trained on FastHTML?

Re: Claude 4

#343
I got Claude 4 Opus to summarize this thread on Hacker News when it had hit 319 comments: https://gist.github.com/simonw/0b9744ae33694a2e03b2169722b06...

Token cost: 22,275 input, 1,309 output = 43.23 cents - https://www.llm-prices.com/#it=22275&ot=1309&ic=15&oc=75&sb=...

Same prompt run against Sonnet 4: https://gist.github.com/simonw/1113278190aaf8baa2088356824bf...

22,275 input, 1,567 output = 9.033 cents https://www.llm-prices.com/#it=22275&ot=1567&ic=3&oc=15&sb=o...

Re: Claude 4

#344

[flagged]

(Caveat on this comment: I'm the COO of OpenRouter. I'm not here to plug my employer; just ran across this and think this suggestion is helpful) Feel free to give OpenRouter a try; part of the value prop is that you purchase credits and they are fungible across whatever models & providers you want. We just got Sonnet 4 live. We have a chatroom on the website, that simply uses the API under the covers (and deducts cre…

Just wanted to provide some (hopefully) helpful feedback from a potential customer that likely would have been, but bounced away due to ambiguity around pricing.

It's too hard to find out what markup y'all charge on top of the APIs. I understand it varies based on the model, but this page (which is what clicking on the "Pricing" link from the website takes you to) https://openrouter.ai/models is way too complicated. My immediate reaction is, "oh shit, this is made for huge enterprises, not for me" followed immediately by "this isn't going to be cheap, I'm not even going to bother." We're building out some AI features in our products so the timing is otherwise pretty good. We're not big fish, but do expect to spending between $3,000 and $5,000 per month once the features hit general availability, so we're not small either. If things go well of course, we'd love to 10x that in the next few years (but time will tell on that one of course).

Re: Claude 4

#345

I can't think of more boring than marginal improvements on coding tasks to be honest. I want GenAI to become better at tasks that I don't want to do, to reduce the unwanted noise from my life. This is when I'll pay for it, not when they found a new way to cheat a bit more the benchmarks. At work I own the development of a tool that is using GenAI, so of course a new better model will be beneficial, especially because…

What if coding is that unwanted task? Also, what are the tasks you are referring to, specifically?

Booking visits to the dentist, hairdresser, any other type of service, renewing my phone or internet subscription at the lowest price, doing administrative tasks, adding the free game of the week on Epic Games Store to my library, finding the right houses to visit, etc.

Basically anything that some startup has tried and failed at uberizing.

Re: Claude 4

#346

An important note not mentioned in this announcement is that Claude 4's training cutoff date is March 2025, which is the latest of any recent model. (Gemini 2.5 has a cutoff of January 2025) https://docs.anthropic.com/en/docs/about-claude/models/overv...

Nice - it might know about Svelte 5 finally...

Re: Claude 4

#347
post #5

Ooh, VS Code integration for Claude Code sounds nice. I do feel like Claude Code works better than the native Cursor agent mode. Edit: How do you install it? Running `/ide` says "Make sure your IDE has the Claude Code extension", where do you get that?

I've been looking for that as well.

Run claude cli from inside of VS Code terminal.

https://news.ycombinator.com/item?id=44064082

Re: Claude 4

#348

Earlier quoted context omitted.

What's with all the models exhibiting sycophancy at the same time? Recently ChatGPT, Gemini 2.5 Pro latest seems more sycophantic, now Claude. Is it deliberate, or a side effect?

IMO, I always read that as a psychological trick to get people more comfortable with it and encourage usage. Who doesn't like a friend who's always encouraging, supportive, and accepting of their ideas?

No. People generally don’t like sycophantic “friends”. Insincere and manipulative.

Re: Claude 4

#349

Earlier quoted context omitted.

Run claude in the terminal in VSCode (or cursor) and it will install automatically!

Doesn't work in Cursor over ssh. I found the VSIX in node_modules/@anthropic-ai/claude-code/vendor and installed it manually.

Found it!

https://news.ycombinator.com/item?id=44064082

Re: Claude 4

#350

Earlier quoted context omitted.

What's with all the models exhibiting sycophancy at the same time? Recently ChatGPT, Gemini 2.5 Pro latest seems more sycophantic, now Claude. Is it deliberate, or a side effect?

IMO, I always read that as a psychological trick to get people more comfortable with it and encourage usage. Who doesn't like a friend who's always encouraging, supportive, and accepting of their ideas?

Me, it immediately makes me think I'm talking to someone fake who's opinion I can't trust at all. If you're always agreeing with me why am I paying you in the first place. We can't be 100% in alignment all the time, just not how brains work, and discussion and disagreement is how you get in alignment. I've worked with contractors who always tell you when they disagree and those who will happily do what you say even if you're obviously wrong in their experience, the disagreable ones always come to a better result. The others are initially more pleasent to deal with till you find out they were just happily going alog with an impossible task.
Post reply on HN