Live data from Hacker News

Claude 4

anthropic.com

71–80 of 1001 posts

Re: Claude 4

#71

I can't think of more boring than marginal improvements on coding tasks to be honest. I want GenAI to become better at tasks that I don't want to do, to reduce the unwanted noise from my life. This is when I'll pay for it, not when they found a new way to cheat a bit more the benchmarks. At work I own the development of a tool that is using GenAI, so of course a new better model will be beneficial, especially because…

yea exactly

Re: Claude 4

#73
post #46

I'm curious what are others priors when reading benchmark scores. Obviously with immense funding at stakes, companies have every incentive to game the benchmarks, and the loss of goodwill from gaming the system doesn't appear to have much consequences. Obviously trying the model for your use cases more and more lets you narrow in on actually utility, but I'm wondering how others interpret reported benchmarks these da…

Well-designed benchmarks have a public sample set and a private testing set. Models are free to train on the public set, but they can't game the benchmark or overfit the samples that way because they're only assessed on performance against examples they haven't seen.

Not all benchmarks are well-designed.

Re: Claude 4

#74
post #13

I'll look at it when this shows up on https://aider.chat/docs/leaderboards/ I feel like keeping up with all the models is a full time job so I just use this instead and hopefully get 90% of the benefit I would by manually testing out every model.

Are these just leetcode exercises? What I would like to see is an independent benchmark based on real tasks in codebases of varying size.

Aider uses a dataset of 500 GitHub issues, so not LeetCode-style work.

Re: Claude 4

#75
Anyone have a link to the actual Anthropic official vscode extension? Struggling to find it.

edit: run `claude` in a vscode terminal and it will get installed. but the actual extension id is `Anthropic.claude-code`

Re: Claude 4

#76
post #70

[flagged]

i agree. i think chatgpt is best model for coding. Coders want to be hardcore and don't want to use the most popular model out there because they don't want do what masses are doing.

Re: Claude 4

#77

How long will the VScode wrapper (cursor, windsurf) survive? Love to try the Claude Code VScode extension if the price is right and purchase-able from China.

what do you mean purchasable from china? As in you are based in china or is there a way to game the tokens pricing

Re: Claude 4

#78
post #51

I've found myself having brand loyalty to Claude. I don't really trust any of the other models with coding, the only one I even let close to my work is Claude. And this is after trying most of them. Looking forward to trying 4.

I've been initially fascinated by Claude, but then I found myself drawn to Deepseek. My use case is different though, I want someone to talk to.

I also use DeepSeek R1 as a daily driver. Combined with Qwen3 when I need better tool usage.

Now that both Google and Claude are out, I expect to see DeepSeek R2 released very soon. It would be funny to watch an actual open source model getting close to the commercial competition.

Re: Claude 4

#79

  > Finally, we've introduced thinking summaries for Claude 4 models that use a smaller model to condense lengthy thought processes. This summarization is only needed about 5% of the time—most thought processes are short enough to display in full.
This is not better for the user. No users want this. If you're doing this to prevent competitors training on your thought traces then fine. But if you really believe this is what users want, you need to reconsider.
Post reply on HN