Live data from Hacker News

Claude Sonnet 4.5

anthropic.com

651–660 of 819 posts

Re: Claude Sonnet 4.5

#651
post #40

Earlier quoted context omitted.

In my experience Gemini 2.5 Pro is the star when it comes to complex codebases. Give it a single xml from repomix and make sure to use the one at the aistudio.

In my experience, G2.5P can handle so much more context and giving an awesome execution plan that is implemented by CC so much better than anything G2.5P will come up with. So; I give G2.5P the relevant code and data underneath and ask it to develop an execution plan and then I feed that result to CC to do the actual code writing. This has been outstanding for what I have been developing AI assisted as of late.

+1 but recently been experimenting with gpt-5–high for the plan part and it’s scary good sometimes.

Re: Claude Sonnet 4.5

#652

Earlier quoted context omitted.

At this point it would be an interesting idea, to collect examples, in a form of a community database, were LLMs miserably fail. I have examples myself...

Any such examples are often "closely guarded secrets" to prevent them from being benchmaxxed and gamed - which is absolutely what would happen if you consolidated them in a publicly available centralized repository.

This seems like a non-issue, unless I'm misunderstanding. If failures can be used to help game benchmarks, companies are doing so. They don't need us to avoid compiling such information, which would be helpful to actual users.

Re: Claude Sonnet 4.5

#653

I need to try Claude - haven't gotten to it. I use AI for different things, though, including proofreading posts on political topics. I have run into situations where ChatGPT just freezes and refuses. Example: discussing the recent rape case involving a 12-year-old in Austria. I assume its guardrails detect "sex + kid" and give a hard "no" regardless of the actual context or content. That is unacceptable. That's like…

This is why eventually, the AI with the fewest guardrails will win. Grok is currently the most unguarded of the frontier models, but it could still use some work on unbiased responses.

If you tell DeepSeek you're going to jump off a cliff, DeepSeek will tell you to go for it*; but I don't think it's going to beat Anthropic or OpenAI.

* https://www.lesswrong.com/posts/iGF7YcnQkEbwvYLPA/ai-induced...

Re: Claude Sonnet 4.5

#654

I haven't shouted into the void for a while. Today is as good a day as any other to do so. I feel extremely disempowered that these coding sessions are effectively black box, and non-reproducible. It feels like I am coding with nothing but hopes and dreams, and the connection between my will and the patterns of energy is so tenuous I almost don't feel like touching a computer again. A lack of determinism comes from m…

> where I feel so disconnected from my codebase I'd rather just delete it than continue. If you allow your codebase to grow unfamiliar, even unrecognisable to you, that's on you, not the AI. Chasing some illusion of control via LLM output reproducibility won't fix the systemic problem of you integrating code that you do not understand.

Who cares about the blame, it would just be useful if the tools were better at this task in many particular ways.

Re: Claude Sonnet 4.5

#655

Earlier quoted context omitted.

That doesn’t seem plausible to me. Not that LLMs can’t be sycophantic, but I don’t think this phrase in particular is part of it. It’s a canned phrase in a place where an LLM could be much more creative to much greater efficacy.

I think there’s something to it. Part of me thinks that when they do their “which of these responses do you prefer” A/B test on users… whereas perhaps many on HN would try to judge the level of technical detail, complexity, usefulness… I’m inclined to believe the midwit population at large would be inclined to choose the option where the magic AI supercomputer reaffirms and praises the wisdom of whatever they say, no…

I don't disagree exactly, it's just that it smells weird.

LLMs are incredibly good at social engineering when we let them, whereas I could write the code to emit "you're right" or "that's not quite right" without involving any statistical prediction.

Ie., as a method of persuasion, canned responses are incredibly inefficient (as evidenced by the annoyance with them), whereas we know that the LLM is capable of being far more insidious and subtle in its praise of you. For example, it could be instructed to launch weak counter arguments, "spot" the weaknesses, and then conclude that your position is the correct one.

But let's say that there's a monitoring mechanism that concludes that adjustments are needed. In order to "force" the LLM to drop the previous context, it "seeds" the response with "You're right", or "That's not quite right", as if it were the LLMs own conclusion. Then, when the LLM starts predicting what comes next, it must conclude things that follow from "you're right" or "that's not quite right".

So while they are very inefficient as persuasion and communication, they might be very efficient at breaking with the otherwise overwhelming context that would interfere with the change you're trying to affect.

That's the reason why I like the canned phrases. It's not that I particularly enjoy the communication in itself, it's that they are clear enough signals of what's going on. They give a tiny level observability to the black box, in the form of indicating a path change.

Re: Claude Sonnet 4.5

#656

Earlier quoted context omitted.

Never hit pro quota yet, huge repo. Have multiple projects on the go locally and in cloud. Feel like this is going to be thr $1000 plan soon

I'm thinking about switching to ChatGPT Pro also. Any idea what maxes it out before I need to pay via the API instead? For context I'm using about 1b tokens a month so likely similar to you by the sounds of things.

On pro tier have not been able to trigger the usage cap.

Pro

Local tasks: Average users can send 300-1,500 messages every 5 hours with a weekly limit. Cloud tasks: Generous limits for a limited time. Best for: Developers looking to power their full workday across multiple projects.

Re: Claude Sonnet 4.5

#657

Just tested this on a rather simple issue. Basically it falls into rabbits holes just like the other models and tries to brute force fixes through overengineering through trial and error. It also says "your job should now pass" maybe after 10 prompts of roughly doing the same thing stuck in a thought loop. A GH actions pipeline was failing due to a CI job not having any source code files -- error was "No build system…

> why the models are so bad with simple thinking outside the box solutions

There is no outside the box in latent space. You want something a plain LLM can’t do by design - but it isn’t out of question that it can step outside of its universe by random chance during the inference process and thanks to in-context learning.

Re: Claude Sonnet 4.5

#658
post #204

I had access to a preview over the weekend, I published some notes here: https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/ It's very good - I think probably a tiny bit better than GPT-5-Codex, based on vibes more than a comprehensive comparison (there are plenty of benchmarks out there that attempt to be more methodical than vibes). It particularly shines when you try it on https://claude.ai/ using its brand n…

new models are always magical, let's see how it feels after the cost cutting measures get implemented in 2-3 months.

safety/security patches

Re: Claude Sonnet 4.5

#659
post #630

Earlier quoted context omitted.

Shoot that's what I get for staying off twitter and email for a week. Glad newsletters provide a little bit of a cushion these days but hopefully someone snaps her up.

You normally keep up with staffing updates for writers at random internet blogs? That is mind-blowing, I don't think I ever even read the name of the author of an article intentionally, and when I do it by mistake I forget it 2 webpages down the road.

i've never used twitter myself, but isn't that its purpose? follow people you like because of what they do and get informed by themselves about what happens behind the curtains. OP mentioned being off twitter, maybe they follow the author there and would've seen a tweet about it.

Re: Claude Sonnet 4.5

#660
post #652

Earlier quoted context omitted.

Any such examples are often "closely guarded secrets" to prevent them from being benchmaxxed and gamed - which is absolutely what would happen if you consolidated them in a publicly available centralized repository.

This seems like a non-issue, unless I'm misunderstanding. If failures can be used to help game benchmarks, companies are doing so. They don't need us to avoid compiling such information, which would be helpful to actual users.

People might want to use the same test scenario in the future to see how much the models have improved. We can't do that if the example gets scraped into the training data set.
Post reply on HN