I can't think of more boring than marginal improvements on coding tasks to be honest. I want GenAI to become better at tasks that I don't want to do, to reduce the unwanted noise from my life. This is when I'll pay for it, not when they found a new way to cheat a bit more the benchmarks. At work I own the development of a tool that is using GenAI, so of course a new better model will be beneficial, especially because…
Claude 4
71–80 of 1001 posts
Re: Claude 4
#72Re: Claude 4
#73I'm curious what are others priors when reading benchmark scores. Obviously with immense funding at stakes, companies have every incentive to game the benchmarks, and the loss of goodwill from gaming the system doesn't appear to have much consequences. Obviously trying the model for your use cases more and more lets you narrow in on actually utility, but I'm wondering how others interpret reported benchmarks these da…
Not all benchmarks are well-designed.
Re: Claude 4
#74I'll look at it when this shows up on https://aider.chat/docs/leaderboards/ I feel like keeping up with all the models is a full time job so I just use this instead and hopefully get 90% of the benefit I would by manually testing out every model.
Are these just leetcode exercises? What I would like to see is an independent benchmark based on real tasks in codebases of varying size.
Re: Claude 4
#75edit: run `claude` in a vscode terminal and it will get installed. but the actual extension id is `Anthropic.claude-code`
Re: Claude 4
#76[flagged]
Re: Claude 4
#77How long will the VScode wrapper (cursor, windsurf) survive? Love to try the Claude Code VScode extension if the price is right and purchase-able from China.
Re: Claude 4
#78I've found myself having brand loyalty to Claude. I don't really trust any of the other models with coding, the only one I even let close to my work is Claude. And this is after trying most of them. Looking forward to trying 4.
I've been initially fascinated by Claude, but then I found myself drawn to Deepseek. My use case is different though, I want someone to talk to.
Now that both Google and Claude are out, I expect to see DeepSeek R2 released very soon. It would be funny to watch an actual open source model getting close to the commercial competition.
Re: Claude 4
#79 > Finally, we've introduced thinking summaries for Claude 4 models that use a smaller model to condense lengthy thought processes. This summarization is only needed about 5% of the time—most thought processes are short enough to display in full.
This is not better for the user. No users want this. If you're doing this to prevent competitors training on your thought traces then fine. But if you really believe this is what users want, you need to reconsider.