Live data from Hacker News

Claude 3.7 Sonnet and Claude Code

anthropic.com

681–690 of 1001 posts

Re: Claude 3.7 Sonnet and Claude Code

#681

You can get your HN profile analyzed by it and it's pretty funny :) https://hn-wrapped.kadoa.com/ I'm using this to test the humor of new models.

>You talk about Amiga computers so much that I'm pretty sure your brain still runs on Kickstart ROM and requires a floppy disk to boot up in the morning.

excuse me, we boot from compact flash these days

>Your comments about modern tech are so critical that I'm convinced you judge new programming languages based on how well they'd run on a Commodore 64.

ouch

Re: Claude 3.7 Sonnet and Claude Code

#682

Earlier quoted context omitted.

Using up to 32k thinking tokens, Sonnet 3.7 set SOTA with a 64.9% score. 65% Sonnet 3.7, 32k thinking 64% R1+Sonnet 3.5 62% o1 high 60% Sonnet 3.7, no thinking 60% o3-mini high 57% R1 52% Sonnet 3.5

How does it stack up against Grok3? I've seen some discussion that Grok3 is good for coding.

It isn't available over api yet, as far as I know. So it can't be really tested independently.

Re: Claude 3.7 Sonnet and Claude Code

#683

Earlier quoted context omitted.

That's a file context problem because you use cursor or cline or some other crap context maker. Try Clood. Unless "anthropic high usage" which I just watch the incident reports I one shot features regularly. At a high skill level. Not front end. Back end c# in a small but great framework that has poor documentation. Not just endpoints but full on task queues. So really, it's a context problem. You're just not laser f…

Wtf is “clood”?

This feels like a technobabble troll. The whole thing is incoherent.

Re: Claude 3.7 Sonnet and Claude Code

#684

Earlier quoted context omitted.

I'm surprised that Gemini 2.0 is first now. I remember that Google models were under performing on kagi benchmarks.

Having your own hardware to run LLMs will pay dividends. Despite getting off on the wrong foot, I still believe Google is best positioned to run away with the AI lead, solely because they are not beholden to Nvidia and not stuck with a 3rd party cloud provider. They are the only AI team that is top to bottom in-house.

We should still wait around to see if Huawei is able to perfect its Ascend series for training and inferencing SOTA models.

Re: Claude 3.7 Sonnet and Claude Code

#685
post #91

Hi everyone! Boris from the Claude Code team here. @eschluntz, @catherinewu, @wolffiex, @bdr and I will be around for the next hour or so and we'll do our best to answer your questions about the product.

What do I need to do to get unbanned? I have filled in the provided Google Docs form 3-4 times to no avail. I got banned almost immediately after joining. My best guess is that I got banned because I used a VPN. https://news.ycombinator.com/item?id=40808815

Re: Claude 3.7 Sonnet and Claude Code

#686

Claude 3.7 Sonnet scored 60.4% on the aider polyglot leaderboard [0], WITHOUT USING THINKING. Tied for 3rd place with o3-mini-high. Sonnet 3.7 has the highest non-thinking score, taking that title from Sonnet 3.5. Aider 0.75.0 is out with support for 3.7 Sonnet [1]. Thinking support and thinking benchmark results coming soon. [0] https://aider.chat/docs/leaderboards/ [1] https://aider.chat/HISTORY.html#aider-v0750

Using up to 32k thinking tokens, Sonnet 3.7 set SOTA with a 64.9% score. 65% Sonnet 3.7, 32k thinking 64% R1+Sonnet 3.5 62% o1 high 60% Sonnet 3.7, no thinking 60% o3-mini high 57% R1 52% Sonnet 3.5

It's clear that progress is incremental at this point. At the same time Anthropic and OpenAI are bleeding money.

It's unclear to me how they'll shift to making money while providing almost no enhanced value.

Re: Claude 3.7 Sonnet and Claude Code

#687
post #638

Earlier quoted context omitted.

I think an argument could be reasonably made that the app layer is the only moat. It’s more likely Anthropic eventually has to acquire Cursor to cement a position here than they out-compete it. Where, why, what brand and what product customers swipe their credit cards for matters — a lot.

Cursor has no models, they dont even have an editor its just vscode

And Typescript simply doesn't work for me. I have tried uninstalling extensions. It is always "Initializing". I reload windows, etc. It eventually might get there, I can't tell what's going on. At the moment, AI is not worth the trade-off of no Typescript support.

Re: Claude 3.7 Sonnet and Claude Code

#688

Very good, Code is extremely nice but as others have said, if you let it go on its own it burns through your money pretty fast. I've made it build a web scraper from scratch, figuring out the "API" of a website using a project from github in another language to get some hints, and while in the end everything was working, I've seen 100k+ tokens being sent too frequently for apparently simple requests, something feels…

It probably makes sense to continue using third party tools such as aider, for now. Anthropic doesn't have a lot of incentives to reduce token usage.

Re: Claude 3.7 Sonnet and Claude Code

#689
post #664
post #655

Earlier quoted context omitted.

But also for $36.83 compared to DeepSeek R1 + claude-3-5 it's $13.29 and for latter "Percent using correct edit format" is 100% vs 97.8% for 3.7. edit: would be interesting to see how combo DeepSeek R1 + claude-3-7 performs.

is there any public info on why such DeepSeek R1 + claude-3-5 combo worked better than using a single model?

Sonnet 3.5 is the best non-Chain-of-Thought code-authoring model. When paired with R1's CoT output, Sonnet 3.5 performs even better - outperforming vanilla R1 (and eveything else), which suggests Sonnet is better than R1 at utilizing R1's CoT.

It's scenario where the result is greater than the sum of it's parts

Re: Claude 3.7 Sonnet and Claude Code

#690
post #479

Earlier quoted context omitted.

That doesn't answer his very specific question.

This thread is giving me a flashback to Stack Overflow

I think I might have just misunderstood his question then. I assumed the limitations come from the subscription plan
Post reply on HN