Live data from Hacker News

Claude 3.7 Sonnet and Claude Code

anthropic.com

911–920 of 1001 posts

Re: Claude 3.7 Sonnet and Claude Code

#911
post #897

I got this working with my LLM tool (new plugin version: llm-anthropic 0.14) and figured out a bunch of things about the model in the process. My detailed notes are here: https://simonwillison.net/2025/Feb/25/llm-anthropic-014/ One of the most exciting new capabilities is that this model has a 120,000 token output limit - up from just 8,000 for the previous Claude 3.5 Sonnet model and way higher than any other model…

No shade against Sonnet 3.7, but I don't think it's accurate to say way higher than any other model in the space. o1 and o3-mini go up to 100,000 output tokens. https://platform.openai.com/docs/models#o1

Huh, good call thanks - I've updated my post with a correction.

Re: Claude 3.7 Sonnet and Claude Code

#912
post #905

Earlier quoted context omitted.

$1.8

I have a very long request to do like this, did you use a specific CLI tool ? (Thank you in advance)

I used my own CLI tool LLM, which can handle these long requests in streaming mode (Anthropic won't let you do a non-streaming request for long output replies like this).

  uv tool install llm
  llm install llm-anthropic
  llm keys set anthropic
  # paste in API key
  llm -m claude-3.7-sonnet -o thinking 1 'your prompt goes here'

Re: Claude 3.7 Sonnet and Claude Code

#913

This AI race is happening so fast. Seems like it to me anyway. As a software developer/engineer I am worried about my job prospects.. time will tell. I am wondering what will happen to the west coast housing bubbles once software engineers lose their high price tags. I guess the next wave of knowledge workers will move in and take their place?

I think if models improve (but we don't get a full singularity) then jobs will increase. e.g. if software is 5x less cost to make, demand will go up more than 5x as supply is highly limited now. Lots of companies want better software but it costs too much. That will create more jobs. They'll be more product management and human interaction and edge case testing and less typing. Although I think there'll be a bunch of…

the thing is that cost won't go down by 5x but much more.

once the ai gets smart enough that it only requires an intern to make the prompt and solve the few mistakes, development cost will be worth nothing.

there is only so much demand for software development.

Re: Claude 3.7 Sonnet and Claude Code

#914
post #756

I am not sure how good these Exercism tasks are for measuring how good at a model with coding. My experience is that these models could write a simple function and get it right if it does not require any out of the box thinking (so essentially offloading the boilerplate part of coding). When it comes to think creatively and have a much better solution to a specific task that would require the think 2-3 steps ahead th…

I think many of the "AI can do coding" narratives don't see what coding means in real situations. It's finding out why "jbdoe1337" added this large if/else around the entire function body back in 2016 - it seems important business logic, but the commit just says "updated code". And how the h*ll this interaction between the conf.ini files, the conf/something.json and the ENV vars works. Why sometimes the ENV var overr…

Systems built from scratch with AI won't have these limitations, because only the model will ever see the code. It will implement a spec that's written in English or another human language.

When the business requirements change, the spec will change. When that happens, the system will either modify its previously-written code or regenerate it from the ground up. Which strategy it chooses won't be especially interesting or important.

The process of maintaining the English-language spec will still require great care and precision. It will be called "programming," or perhaps "coding."

A few graybearded gurus will insist on examining the underlying C or Javascript or Python or Rust or whatever the model generates, the way they peer at compiler-generated assembly code now. Occasionally this capability will be important, even vital. But not usually. The situations where it's necessary will become less common over time.

Re: Claude 3.7 Sonnet and Claude Code

#915

Claude 3.7 Sonnet scored 60.4% on the aider polyglot leaderboard [0], WITHOUT USING THINKING. Tied for 3rd place with o3-mini-high. Sonnet 3.7 has the highest non-thinking score, taking that title from Sonnet 3.5. Aider 0.75.0 is out with support for 3.7 Sonnet [1]. Thinking support and thinking benchmark results coming soon. [0] https://aider.chat/docs/leaderboards/ [1] https://aider.chat/HISTORY.html#aider-v0750

Hi Paul, been following the aider project for about a year now to develop an understanding of how to build SWE agents.

I was at the AI Engineering Summit in NYC last week and met an (extremely senior) staff ai engineer doing somewhat unbelievable things with aider. Shocking things tbh.

Is there a good way to share stories about real-world aider projects like this with you directly (if I can get approval from him)? Not sure posting on public forum is appropriate but I think you would be really interested to hear how people are using this tool at the edge.

Re: Claude 3.7 Sonnet and Claude Code

#916

Claude 3.7 Sonnet scored 60.4% on the aider polyglot leaderboard [0], WITHOUT USING THINKING. Tied for 3rd place with o3-mini-high. Sonnet 3.7 has the highest non-thinking score, taking that title from Sonnet 3.5. Aider 0.75.0 is out with support for 3.7 Sonnet [1]. Thinking support and thinking benchmark results coming soon. [0] https://aider.chat/docs/leaderboards/ [1] https://aider.chat/HISTORY.html#aider-v0750

Is aider still relevant vs. Claude Code?

Yes. Absolutely it is. For different workloads it is an insanely effective tool.

Re: Claude 3.7 Sonnet and Claude Code

#917

You can get your HN profile analyzed by it and it's pretty funny :) https://hn-wrapped.kadoa.com/ I'm using this to test the humor of new models.

>You'll create a browser extension that automatically bypasses paywalls and archives important articles - because apparently saving democracy shouldn't cost $12.99/month

>Your archive.is links will become so legendary that dang will create a special 'Paywall Slayer' badge just for you

>You've shared so many archive.is links that the Internet Archive is considering naming you their unofficial spokesperson - or sending you a cease and desist letter.

>Your economic predictions are so consistently apocalyptic that gold dealers use your comment history as their marketing strategy.

Really sums it up!

Re: Claude 3.7 Sonnet and Claude Code

#918

You can get your HN profile analyzed by it and it's pretty funny :) https://hn-wrapped.kadoa.com/ I'm using this to test the humor of new models.

> Your salary is so low even your legacy code feels sorry for you. > You're the only person on HN who thinks $800/month is a salary and not a cloud computing bill. ouch

>You're the only person on HN who thinks $800/month is a salary and not a cloud computing bill.

Now that is funny!

Re: Claude 3.7 Sonnet and Claude Code

#919

You can get your HN profile analyzed by it and it's pretty funny :) https://hn-wrapped.kadoa.com/ I'm using this to test the humor of new models.

Your comments about suburban missile defense systems have the FBI agent monitoring your internet connection seriously questioning their career choices. You've spent so much time explaining why manufacturing is complex that you could have just built your own CRT factory by now. You claim to be skeptical of AI hype, yet you've indexed more documentation with Cursor than most people have read in their lifetime. Surprisi…

Yeah it seems like it doesn't go back very far, which is understandable.

Re: Claude 3.7 Sonnet and Claude Code

#920

Earlier quoted context omitted.

How does it stack up against Grok3? I've seen some discussion that Grok3 is good for coding.

Pro tip: It's hard to trust Twitter for opinions on Grok. The thumb is very clearly on the scale. I have personally seen very few positive opinions of Grok outside of Twitter.

I agree with you, and I hate to say this, but I saw them on LinkedIn. One purportedly used the same prompts to make a "pacman like" game and the results from Grok3 were at least better, assuming the post is on the up and up, better looking than o3-mini-high.
Post reply on HN