Live data from Hacker News

Claude 3.7 Sonnet and Claude Code

anthropic.com

81–90 of 1001 posts

Re: Claude 3.7 Sonnet and Claude Code

#83

They don't say this, but from querying it, they also seem to have updated the knowledge cutoff from April 2024 ("3.6") to October 2024 (3.7)

It's in the Model Card: https://assets.anthropic.com/m/785e231869ea8b3b/original/cla...

>Claude 3.7 Sonnet is trained on a proprietary mix of publicly available information on the Internet as of November 2024

Re: Claude 3.7 Sonnet and Claude Code

#85
post #27

Earlier quoted context omitted.

It does seem like it will be very, very hard for the companies training their own models to recoup their investment when the capabilities of open-weight models catch up so quickly - general purpose LLMs just seem destined to be a cheap commodity.

Well, the companies releasing open weights also need to recoup their investments at some point, they can't coast on VC hype forever. Huge models don't grow on trees.

Or, like Meta, they make their money elsewhere and just seem interested in wrecking the economics of LLMs. As soon as an open-weight model is released, it basically sets a global floor that says "Models with similar or worse performance effectively have zero value," and that floor has been rising incredibly quickly. I'd be surprised if the vast, vast majority of queries ChatGPT gets couldn't get equivalently good results from llama3/deepseek/qwen/mistral models, even for those paying for the pro versions.

Re: Claude 3.7 Sonnet and Claude Code

#88
post #3

Anthropic doubling down on code makes sense, that has been their strong suit compared to all other models Curious how their Devin competitor will pan out given Devin's challenges

It's their strong suit no doubt, but sometimes I wish the chat would not be so eager to code. It often throws code at me when I just want a conceptual or high level answer. So often that I routinely tell it not to.

> I just want a conceptual or high level answer

I've found claude to be very receptive to precise instructions. If I ask for "let's first discuss the architecture" it never produces code. Aider also has this feature with /architect

Re: Claude 3.7 Sonnet and Claude Code

#89

Earlier quoted context omitted.

It's their strong suit no doubt, but sometimes I wish the chat would not be so eager to code. It often throws code at me when I just want a conceptual or high level answer. So often that I routinely tell it not to.

I complain about this all the time, despite me saying "ask me questions before you code" or all these other instructions to code less, it is SO eager to code. I am hoping their 3.7 reasoning follows these instructions better

We should remember 3.5 was trained in an era when ChatGPT would routinely refuse to code at all and architected in an era when system prompts were not necessarily very effective. I bet this will improve, especially now that Claude has its own coding and arch cli tool.

Re: Claude 3.7 Sonnet and Claude Code

#90

To me the biggest surprise was seeking grok dominate in all of their published benchmarks. I haven’t seen any benchmarks of it yet (which I take with a giant heap of salt), but it’s still interesting nevertheless. I’m rooting for Anthropic.

Yeah, putting it on the opposite side of that comparison chart was a sleezy but likely effective move.
Post reply on HN