Live data from Hacker News

Claude 3.7 Sonnet and Claude Code

anthropic.com

811–820 of 1001 posts

Re: Claude 3.7 Sonnet and Claude Code

#811
post #790
post #774

Earlier quoted context omitted.

> Phrased this way without any help, all but the thinking models get it wrong I C&P'd it into Claude 3.7 with thinking, and it gave the correct answer (which I'm pretty sure is #2). Including the CoT, where it actually does math (which I haven't checked), and final response. # THINKING Let's analyze the two options. Option 1: Add cold milk immediately, then let it sit for 2 mins. Option 2: Let it sit for 2 mins, then…

Perhaps use pastebin for synthetic content next time?

Thanks for the heads-up; I was pretty confused why I was getting downvoted, as it seemed like "Here's a counterexample to your claim" is pretty on-topic.

Unfortunately I only noticed it after the window to edit the comment was closed. If the first person to downvote me had instead suggested I use a pastebin, I might have been able to make the conversation more agreeable to people.

Re: Claude 3.7 Sonnet and Claude Code

#812
post #8
post #3

Anthropic doubling down on code makes sense, that has been their strong suit compared to all other models Curious how their Devin competitor will pan out given Devin's challenges

I thought the same thing, I have 3 really hard problems that Claude (or any model) hasn’t been able to solve so far and I’m really excited to try them today

Did it work?

Re: Claude 3.7 Sonnet and Claude Code

#813

Earlier quoted context omitted.

** Roast *** * You've spent more time talking about your Carnatic raga detector than actually building it – at this rate, LLMs will be composing ragas before your detector can identify them. * You bought a 7950X processor but can't figure out what to do with it – the computing equivalent of buying a Ferrari to drive to the grocery store once a week. * You're so concerned about work-life balance that you took a sabbat…

That sabbatical one is savage.

Funnily enough, I'm putting the 7950X to some use in the Carnatic Raga detector project since a lot of audio operations are heavy on CPU. But that last one nearly killed me. I'll have to go to Gemini or GPT for some therapy after that one.

Re: Claude 3.7 Sonnet and Claude Code

#814

Kagi LLM benchmark updated with general purpose and thinking mode for Sonnet 3.7. https://help.kagi.com/kagi/ai/llm-benchmark.html Appears to be second most capable general purpose LLM we tried (second to gemini 2.0 pro, in front of gpt-4o). Less impressive in thinking mode, about at the same level as o1-mini and o3-mini (with 8192 token thinking budget). Overall a very nice update, you get higher quality and higher…

I thought o3-mini was o1-mini. OpenAI's naming gets confusing.

Re: Claude 3.7 Sonnet and Claude Code

#815

Earlier quoted context omitted.

Using up to 32k thinking tokens, Sonnet 3.7 set SOTA with a 64.9% score. 65% Sonnet 3.7, 32k thinking 64% R1+Sonnet 3.5 62% o1 high 60% Sonnet 3.7, no thinking 60% o3-mini high 57% R1 52% Sonnet 3.5

How does it stack up against Grok3? I've seen some discussion that Grok3 is good for coding.

Pro tip: It's hard to trust Twitter for opinions on Grok. The thumb is very clearly on the scale. I have personally seen very few positive opinions of Grok outside of Twitter.

Re: Claude 3.7 Sonnet and Claude Code

#816
post #114

I'm about 50kloc into a project making a react native app / golang backend for recipes with grocery lists, collaborative editing, household sharing, so a complex data model and runtime. Purely from the experiment of "what's it like to build with AI, no lines of code directly written, just directing the AI." As I go through features, I'm comparing a matrix of Cursor, Cline, and Roo, with the various models. While I'm…

"no lines of code directly written, just directing the AI" /skeptical face. Without fail, every. single. person. I've met who says that, actually means "except for the code that I write", or "except for how I link the code it build together by hand". If you are 50kloc in to a large complex project that you have literally written none of, and have, eg. used cursor to generate the code without any assistance... well, y…

That's the point of the experiment I'm doing, what it takes to get these things to be able to generate all the code, and I'm just directing.

I literally have not written a line of code. The AI agent configures the build systems. It executes the `go install` command. It configures the infrastructure via terraform.

It takes a lot of reading of the code that's generated to see what I agree with or not, and redirecting refactorings. Understanding how to describe problem statements that are translated into design docs that are translated into task lists. It's still a lot of knowledge work on how to build software. But now I can do the coding that might have taken a day from those plans in 20 minutes.

Regarding startups, there's nothing here I'm doing that isn't just learning the tools of agentic coding. The business here might be advising people on how to do it themselves.

Re: Claude 3.7 Sonnet and Claude Code

#818

You can get your HN profile analyzed by it and it's pretty funny :) https://hn-wrapped.kadoa.com/ I'm using this to test the humor of new models.

Huh, interesting what it focused on.

> You've cited LessWrong so many times that Eliezer Yudkowsky is considering charging you royalties for intellectual property use. > Your comments have more 'bits of evidence' and 'probability updates' than most scientific papers. Have you considered that sometimes people just want to chat without Bayesian analysis? > You spend so much time trying to bring nuance to political discussions on HN that you could have single-handedly solved AI alignment by now.

Re: Claude 3.7 Sonnet and Claude Code

#819

Earlier quoted context omitted.

How does it stack up against Grok3? I've seen some discussion that Grok3 is good for coding.

Pro tip: It's hard to trust Twitter for opinions on Grok. The thumb is very clearly on the scale. I have personally seen very few positive opinions of Grok outside of Twitter.

I thought Grok 2 was pretty bad, but Grok 3 is actually quite good. I'm mostly impressed by the speed of answering. But Claude is still the king of code.

Re: Claude 3.7 Sonnet and Claude Code

#820

Earlier quoted context omitted.

It has the potential to effect a lot more than just SV/The West Coast - in fact SV may be one of the only areas who have some silver lining with AI development. I think these models have a chance to disrupt employment in the industry globally. Ironically it may be only SWE's and a few other industries (writing, graphic design, etc) that truly change. You can see they and other AI labs are targeting SWEs in particular…

What do you even do then as a student? I've asked this dozens of times with zero practical answers at all. Frankly I've become entirely numb to it all.

I'm sure lots of potential students / bootcampers are now not going into programming (or if they are, the smart ones try to go into niches like A.I and skip web/backend/android altogether). This will work against the numbers of jobs being reduced by A.I. It will take a few years though to play out , but at some point we will see smaller amounts of people trying to get into the field and applying for jobs, certainly for junior positions. We've already had ~ 2 bad years, a couple more like this will really dry out the numbers of newcomers. Less people coming in (than otherwise would have) means for every person who retires / leaves the industry there are less people to take his place. This situation is quite complex with lots of parameters that work in different directions so it's very early to try to get some kind of read on where this is going.

As a new career I'd probably not choose SWE now. But if you've done 10 years already I'd ride it out, there is a good chance most of us will remain employed for many years to come.

Post reply on HN