Live data from Hacker News

Claude Sonnet 4.5

anthropic.com

621–630 of 819 posts

Re: Claude Sonnet 4.5

#621
post #96
post #57

Earlier quoted context omitted.

Also getting a perfect score on AIME (math) is pretty cool. Tongue in cheek: if we progress linearly from here software engineering as defined by SWE bench is solved in 23 months.

Just a few months ago people were still talking about exponential progress. The fact that we’re already going for just linear progress is not a good sign

https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com...

We are still at 7mo doubling time on METR task duration. If anything the rate is increasing if you bias to more recent measurements.

Re: Claude Sonnet 4.5

#622

Earlier quoted context omitted.

Very interesting observation. I haven’t written a function by hand in 18 months.

Have you built anything in 18 months? I keep asking to see these apps that people supposedly vibe coded in a weekend but when I ask them to share it, nothing.

https://www.gabrieluribe.me

I have some publicly accessible projects there.

Re: Claude Sonnet 4.5

#623

Earlier quoted context omitted.

It gets annoying because A) it so quickly dismisses its own logic and conclusion from less than two minutes ago (extreme confidence with minimal conviction), and B) it fucks up the second time too (sometimes in the same way!) about 33% of the time.

Gemini 2.5 Pro seems to have a tic where after an initial failed task, it then starts asserting escalating levels of confidence for each subsequent attempt. Like it's ever conscious of its failure lingering in its context and feels the need to over compensate as a form of reassuring both the user and itself that it's not going to immediately faceplant again.

ChatGPT does the same thing, to the point that after several rounds of pointing out errors or hallucinations it will say things like “Ok, you’re right. No more foolish mistakes. This is it, for all the marbles. Here is an assured, triple-checked, 100% error-free, working script, with no chance of failure.”

Which fails in pretty much the exact same way it did before.

Once ChatGPT hits that supremely confident “Ok nothing was working because I was being an idiot but now I’m not” type of dialogue, I know it’s time to just start a new chat. There’s no pulling it out of “spinning the tires while gaslighting” mode.

I’ve even had it go as far as outputting a zip file with an empty .txt that supposedly contained the solution to a certain problem it was having issues with.

Re: Claude Sonnet 4.5

#624

Earlier quoted context omitted.

I'm not trying to be offensive here, feel the need to indicate that. But that prompt leads me to believe that you're going to get rather 'random' results due to leaving SO much room for interpretation. Also, in my experience, punctuation is important - particularly for pacing and grouping of logical 'parts' of a task and your prompt reads like a run on sentence. Making a lot of assumptions here - but I bet if I were…

I think that is an interesting observation and I generally agree. Your point about prompting quality is very valid and for larger features I always use PRDs that are 5-20x the prompt. The thing is my "experiment" is one that represents a fairly common use case: this feature is actually pretty small and embeds into an pre-existing UI structure - in a larger codebase. GPT-5-Codex allows me to write a pretty quick & dir…

I would think that to truly rank such things, you should run a few tests and look for a clear pattern. It's possible that something promoted claude to take "the easy way" while chatgpt didn't.

Re: Claude Sonnet 4.5

#625
post #390

Earlier quoted context omitted.

Don't be so grim! This will just give you access to not worry about writing clean code as much as you did in the past - you can focus on other parts of the development lifecycle. The skill of writing good quality code is still going to be beneficial, maybe less emphasized on writing side, but critical of shipping good code, even when someone (something) else wrote it.

“Do t worry about the fit and finish in your craftsmanship anymore, just bolt everything together and move on to other woodworking” Is how that argument comes across.

This seems broadly correct? Industrialization was amazing for people's standard of living, but it absolutely meant that the average physical good became detached from their craftsmen's learned and aesthetic preferences.

Re: Claude Sonnet 4.5

#626

To @simonw and all the coding agent and LLM benchmarkers out there: please, always publish the elapsed time for the task to complete successfully! I know this was just a "it works straight in claude.ai" post, but still, nowhere in the transcript there's a timestamp of any kind. Durations seem to be COMPLETELY missing from the LLM coding leaderboards everywhere [1] [2] [3] There's a huge difference in time-to-completi…

This is very relevant to this release. It’s way faster, but also seems lazier and more likely to say something’s done when it isn’t (at least in CC). On net it feels more productive because all the small “more padding” prompts are lightning fast, and the others you can fix.

Re: Claude Sonnet 4.5

#627

Anecdotal evidence. I have a fairly large web application with ~200k LoC. Gave the same prompt to Sonnet 4.5 (Claude Code) and GPT-5-Codex (Codex CLI). "implement a fuzzy search for conversations and reports either when selecting "Go to Conversation" or "Go to Report" and typing the title or when the user types in the title in the main input field, and none of the standard elements match, a search starts with a 2s de…

My first thought was I bet I could get Sonnet to fix it faster because I got something back in 3 minutes instead of 20 minutes. You can prompt a lot of changes with a faster model. I'm new to Claude Code, so generally speaking I have no idea if I'm making sense or not.

Re: Claude Sonnet 4.5

#628

Earlier quoted context omitted.

This is obviously much more than just taking an LLM an letting it run for 30 hours. You have to build a whole environment together with external tool integration and context management and then tune the prompts and perhaps even set up a multi-agent system. I believe that if someone puts a ton of work into this you can have an LLM run for that long and still produce sellable outputs, but let's not pretend like this is…

Well, yes, that's Claude Code. And OpenAI Codex. And Google Gemini CLI. Your average dev can just use those.

Yes but you need to setup quite a bit of tooling to provide feedback loops.

It's one thing to get an llm to do something unattended for long durations, it's a other to give it the means of verification.

For example I'm busy upgrading a 500k LoC rails 1 codebase to rails 8 and built several DSLs that give it proper authorised sessions in a headless browser with basic html parsing tooling so it can "see" what affect it's fixes have. Then you somehow need to also give it a reliable way to keep track of the past and it's own learnings, which sound simple but I have yet to see any tool or model solve it on this scale...will give sonnet 4.5 a try this weekend, but yeah none of the models I tried are able to produce meaningful results over long periods on this upgrade task without good tooling and strong feedback loops

Btw I have upgraded the app and taking it to alpha testing now so it is possible

Re: Claude Sonnet 4.5

#629

Earlier quoted context omitted.

That minutiae was always borderline irrelevant, the skill was always making somebody money, possibly with software. The reality is that more software will be pushed than before, and more of it will need to be overseen by a professional.

> minutiae was always borderline irrelevant, the skill was always making somebody money That's extremely reductive, and a prime example of why everything is enshittified today.

How is enshitification (the gradual degredation of service and products for commercial gain) even related to what's being discussed (the gradual obsoletion of a certain set of skills of an SWE)?

Re: Claude Sonnet 4.5

#630

Earlier quoted context omitted.

Although she is indeed solid as an AI journalist, unfortunately she was recently let go for unknown reasons: https://www.kyliebytes.com/thank-god-i-got-fired/

Shoot that's what I get for staying off twitter and email for a week. Glad newsletters provide a little bit of a cushion these days but hopefully someone snaps her up.

You normally keep up with staffing updates for writers at random internet blogs? That is mind-blowing, I don't think I ever even read the name of the author of an article intentionally, and when I do it by mistake I forget it 2 webpages down the road.
Post reply on HN