Live data from Hacker News

Claude Sonnet 4.5

anthropic.com

451–460 of 819 posts

Re: Claude Sonnet 4.5

#451

Anecdotal evidence. I have a fairly large web application with ~200k LoC. Gave the same prompt to Sonnet 4.5 (Claude Code) and GPT-5-Codex (Codex CLI). "implement a fuzzy search for conversations and reports either when selecting "Go to Conversation" or "Go to Report" and typing the title or when the user types in the title in the main input field, and none of the standard elements match, a search starts with a 2s de…

I think Codex working for 20 mins uninterrupted is actually a strength. It’s not “slow” as critics sometimes say - it’s thorough and autonomous. I can actually walk away and get something else done around the house while it does my work for me.

I swear cc in June/July used to spend a lot more time on tasks and felt more thorough like codex does now. Hard to remember much past the last week in this world though.

Re: Claude Sonnet 4.5

#452
post #380

Does 4.5 still answer everything with "You're absolutely right!" or is it now able to communicate like a real programmer?

It still says "Perfect!" about its own work far too often.

Re: Claude Sonnet 4.5

#454
post #380

Does 4.5 still answer everything with "You're absolutely right!" or is it now able to communicate like a real programmer?

I won’t be satisfied until I get a Linus Torvalds mode. “Your idea is shit because you are so fucking stupid” “Please stop talking, it hurts my GPUs thinking down to your level” “I may seem evil but at least I’m not incompetent”

I'm pretty sure you could get Grok 4 to do that without much trouble.

Re: Claude Sonnet 4.5

#455

To @simonw and all the coding agent and LLM benchmarkers out there: please, always publish the elapsed time for the task to complete successfully! I know this was just a "it works straight in claude.ai" post, but still, nowhere in the transcript there's a timestamp of any kind. Durations seem to be COMPLETELY missing from the LLM coding leaderboards everywhere [1] [2] [3] There's a huge difference in time-to-completi…

That's a good call, I'll try to remember that for next time.

Re: Claude Sonnet 4.5

#457
Claude doesn't know how to calculate realistic minimum voltages for solar arrays w/MPPT chargers. ChatGPT does.

Prompt: "Can I use two strings of four Phono Solar PS440M8GFH solar panels with a EG4 12kPV Hybrid Inverter? I want to make sure that there will not be an issue any time of year. New York upstate."

Claude 4.5: Returns within a few seconds. Does not find the PV panel specs, so it asks me if I want it to search for them. I say yes. Then it finally comes up with: "YES, your configuration is SAFE [...] MPPT range check: Your operating voltage of 131.16V fits comfortably in the 120-500V MPPT operating range".

ChatGPT 5: Returns after 78 seconds. Says: "Hot-weather Vmpp check: Vmpp_string @ STC = 4 × 32.79 = 131 V (inside 120–500 V). Using the panel’s NOCT point (31.17 V each), a typical summer operating point is ~125 V — still OK. But at very hot cell temps (≈70 °C is possible), Vmpp can drop roughly ~13% from STC → ~114 V, which is below the EG4’s 120 V MPPT lower limit. That can cause the tracker to fall out of its optimal range and reduce harvest during peak heat."

ChatGPT used deeper thinking to determine that the lowest possible voltage in the heat would be below the MPPT's minimum operating voltage. It doesn't indicate that in reality it might not charge at all at that point... but it does point out the risk, whereas Claude says everything is fine. I need about 5 back-and-forths with Claude to get it to finally realize its mistake.

Re: Claude Sonnet 4.5

#458

Earlier quoted context omitted.

Why did you have access to a preview?

His "pelican riding a bicycle" tests are now a classic and AI shops are benchmaxxing for it

If they were testing that it'd work more often.

Other things you can ask that they're still clearly not optimizing for are ASCII art and directions between different locations. Complete fabrications 100% of the time.

Re: Claude Sonnet 4.5

#459
post #380

Does 4.5 still answer everything with "You're absolutely right!" or is it now able to communicate like a real programmer?

I won’t be satisfied until I get a Linus Torvalds mode. “Your idea is shit because you are so fucking stupid” “Please stop talking, it hurts my GPUs thinking down to your level” “I may seem evil but at least I’m not incompetent”

I'm still holding out for _Marvin the depressed robot from Hitchhiker's Guide_ mode. "Why does anyone program anything?"

Re: Claude Sonnet 4.5

#460
post #455

To @simonw and all the coding agent and LLM benchmarkers out there: please, always publish the elapsed time for the task to complete successfully! I know this was just a "it works straight in claude.ai" post, but still, nowhere in the transcript there's a timestamp of any kind. Durations seem to be COMPLETELY missing from the LLM coding leaderboards everywhere [1] [2] [3] There's a huge difference in time-to-completi…

That's a good call, I'll try to remember that for next time.

I just wanted to say that I really liked your this comment which just showed professionalism and just learning from your mistakes/improving yourself.

I definitely consider you to be an AI influencer, especially in hackernews communities and so I wanted to say that I see influencers who will double down,triple down on things when in reality, people just wanted to help them in the first place.

I just wanted to say thanks with all of this in mind, also that your generate me a pelican riding a bicycle has been a fun ride and is always going to be interesting, so thanks for that as well. I just wanted to share my gratitude with ya.

Post reply on HN