Live data from Hacker News

Claude 3.7 Sonnet and Claude Code

anthropic.com

731–740 of 1001 posts

Re: Claude 3.7 Sonnet and Claude Code

#731

Earlier quoted context omitted.

Not enough resources to get another bachelors, and a masters is probably practically worthless for a pivot. I would have to throw away the past 10 years of my life, start from scratch, with zero ideas for any real skill-developing projects since I'm not interested at all. Probably a completely non-viable candidate in anything I would choose. Maybe only Robotics would work, and that's probably going to be solved quick…

I think if you believe LLMs can truly generalize and will be able to replace all labor in entire industries and 10x every year, you pretty much should believe in ASI at which point having a job is the least of your problems. if you rule out ASI, then that means progress is going to have to slow. consider that programming has been getting more and more automated continually since 1954. so put yourself in a position wh…

I don't know if I agree with that and as a SWE myself its tempting to think that - it it a form of coping and hope that we will be all in it together.

However rationally I can see where these models are evolving, and it leads me to think the software industry is on its own here at least in the short/medium term. Code and math, and with math you typically need to know enough about the domain know what abstract concept to ask, so that just leaves coding and software development. Even for non technical people they understand the result they want of code.

You can see it in this announcement - it's all about "code, code, code" and how good they are in "code". This is not by accident. The models are becoming more specialised and the techniques used to improve them beyond standard LLM's are not as general to a wide variety of domains.

We engineers think AI automation is about difficulty and intelligence, but that's only partly true. Its also about whether the engineer has the knowledge on what they want to automate, the training data is accessible and vast, and they even know WHAT data is applicable. This combination of both deep domain skills and AI expertise is actually quite rare which is why every AI CEO wants others to go "vertical" - they want others to do that leg work on their platforms. Even if it eventuates it is rare enough that, if they automate, will automate a LOT slower not at the deltas of a new model every few months.

We don't need AGI/ASI to impact the software industry; in my opinion we just need well targeted models that get better at a decent rate. At some point they either hit a wall or surpass people - time will tell BUT they are definitely targeting SWE's at this point.

Re: Claude 3.7 Sonnet and Claude Code

#732

Kagi LLM benchmark updated with general purpose and thinking mode for Sonnet 3.7. https://help.kagi.com/kagi/ai/llm-benchmark.html Appears to be second most capable general purpose LLM we tried (second to gemini 2.0 pro, in front of gpt-4o). Less impressive in thinking mode, about at the same level as o1-mini and o3-mini (with 8192 token thinking budget). Overall a very nice update, you get higher quality and higher…

I see it in Kagi Assistant already and it's not even 24 hours! Nice.

Re: Claude 3.7 Sonnet and Claude Code

#733
Is it just me who get the feeling that Claude 3.7 is worse than 3.5?

I really like 3.5 and can be productive with it, but with Claude 3.7 it can't fix even simple things.

Last night I sat for 30 minutes just to try to get the new model to remove a instructions section from a Next.js page. It was an isolated component on the page named InstructionsComponent. Failed non-stop, didn't matter what I did, it could not do it. 3.5 did it first try, I even mistyped instructions and the model fixed the correct thing anyway.

Re: Claude 3.7 Sonnet and Claude Code

#735
post #729

Ahha, recently my daugher come to me with 3rd grade math problem. "Without rearranging the digits 1 2 3 4 5, insert mathematical operation signs and, if necessary, parentheses between them so that the resulting expression equals 40 and 80. The key is that you can combine digits (like 12+3/45) but you cannot change their order from the original sequence 1,2,3,4,5" Grok3, Claude, Deepseek, Qwen all failed to solve this…

This is what they are expecting 3rd graders to solve in math? Pretty hard for that age?

Re: Claude 3.7 Sonnet and Claude Code

#736
post #706
post #691

Earlier quoted context omitted.

I agree that retrieval can take many forms besides vector search, but do we really want to call it RAG if the model is directing the search using a tool call? That like an important distinction to me and the name "agentic search" makes a lot more sense IMHO.

Yes, I think that's RAG. It's Retrieval Augmented Generation - you're retrieving content to augment the generation. Who cares if you used vector search for the retrieval? The best vector retrieval implementations are already switching to a hybrid between vector and FTS, because it turns out BM25 etc is still a better algorithm for a lot of use-cases. "Agentic search" makes much less sense to me because the term "agen…

I think it depends who "you" is. In classic RAG the search mechanism is preordained, the search is done up front and the results handed to the model pre-baked. I'd interpret "agentic search" as anything where the model has potentially a collection of search tools that it can decide how to use best for a given query, so the search algorithm, the query, and the number of searches are all under its own control.

Re: Claude 3.7 Sonnet and Claude Code

#737

You can get your HN profile analyzed by it and it's pretty funny :) https://hn-wrapped.kadoa.com/ I'm using this to test the humor of new models.

> A cryptography enthusiast who created Coze and spends their days defending proper base encoding practices while reminding everyone about the forgotten 33rd ASCII control character.

The nerd humor was hilariously unexpected.

> Your deep dives into quantum mechanics will lead you to publish a paper reconciling quantum eraser experiments with your cryptographic work, confusing physicists and cryptographers alike.

That is one hell of a Magic 8 Ball.

https://hn-wrapped.kadoa.com/Zamicol

Re: Claude 3.7 Sonnet and Claude Code

#738

Earlier quoted context omitted.

> 225 coding exercises from Exercism Has there been any effort taken to reduce data leakage of this test set? Sounds like these exercises were available on the internet pre-2023, so they'll probably be included in the training data for any modern model, no?

I try not to let perfect be the enemy of good. All benchmarks have limitations. The Exercism problems have proven to be very effective at measuring an LLM's ability to modify existing code. I receive a lot of feedback that the aider benchmarks correlate strongly with people's "vibes" on model coding skill. I agree. The scores have felt quite aligned with my hands-on experience coding with most of the top models over…

That's my perception as well. Most of the time, most of the devs I know, including myself, are not really creating novelty with the code itself, but with the product. (Sometimes even the product is not novel, just a similar enhanced version of existing products)

If the resulting code is not trying to be excessively clever or creative this is actually a good thing in my book.

The novelty and creativity should come from the product itself, especially from the users/customers perspective. Some people are too attached to LLM leaderboards being about novelty. I want reliable results whenever I give the instructions, either be the code, or the specs built into a spec file after throwing some ideas into prompts.

Re: Claude 3.7 Sonnet and Claude Code

#740

Earlier quoted context omitted.

It's a phone number. It's probably been bought / sold a few times already. Unless you're on the level of Edward Snowden, I wouldn't worry about it. But maybe your sense of privacy is more valuable than the outcome you'd get from Claude. That's fine too.

It's my phone number... linked to my Google identity... linked to every submitted user prompt... linked to my source code. There's also been a spate of AI companies rushing to release products and having "oops" moments where they leaked customer chats or whatever. They're not run like a FAANG, they don't have the same security pedigree, and they generally don't have any real guarantee of privacy. So yes, my privacy i…

Just buy a $5 burner phone number. No need to use your real one.
Post reply on HN