Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…
Claude Sonnet 4.5
111–120 of 819 posts
Re: Claude Sonnet 4.5
#112Re: Claude Sonnet 4.5
#113Also as a Max $200 user, feels weird to be paying for an Opus tailored sub when now the standard Max $100 would be preferred since they claim Sonnet is better than Opus.
Hope they have Opus 4.5 coming out soon or next month i'm downgrading.
Re: Claude Sonnet 4.5
#114I used to use cc, but I switched to codex (and it was much better) ... no I guess I have to switch batch to CC, at least to test it
Re: Claude Sonnet 4.5
#115Earlier quoted context omitted.
GPT-5 is like the guy on the baseball team that's really good at hitting home runs but can't do basic shit in the outfield. It also consistently gets into drama with the other agents e.g. the other day when I told it we were switching to claude code for executing changes, after badmouthing claude's entirely reasonable and measured analysis it went ahead and decided to `git reset --hard` even after I twice pushed back…
Please tell me you're joking or at least exaggerating about GPT-5's behavior
To be clear, I don't believe that there was any _intention_ of malice or that the behavior was literally envious in a human sense. Moreso I think they haven't properly aligned GPT-5 to deal with cases like this.
Re: Claude Sonnet 4.5
#116I use AI for different things, though, including proofreading posts on political topics. I have run into situations where ChatGPT just freezes and refuses. Example: discussing the recent rape case involving a 12-year-old in Austria. I assume its guardrails detect "sex + kid" and give a hard "no" regardless of the actual context or content.
That is unacceptable.
That's like your word processor refusing to let you write about sensitive topics. It's a tool, it doesn't get to make that choice.
Re: Claude Sonnet 4.5
#117I just ran this through a simple change I’ve asked Sonnet 4 and Opus 4.1, and it fails too. It’s a simple substitution request where I provide a Lint error that suggests the correct change. All the models fail. I could ask someone with no development experience to do this change and they could. I worry everyone is chasing benchmarks to the detriment of general performance. Or the next token weight for the incorrect c…
More like churning benchmarks... Release new model at max power, get all the benchmark glory, silently reduce model capability in the following weeks, repeat by releasing newer, smarter model.
The only way around this is to never report on the same benchmark versions twice, which they include too many to realistically do every release.
Re: Claude Sonnet 4.5
#118Re: Claude Sonnet 4.5
#119Earlier quoted context omitted.
Because it excercises thinking about a pelican riding a bike (not common) and then describing that using SVG. It's quite nice imho and seems to scale with the power of the LLM model. Sure Simon has some actual reasons though.
> Because it excercises thinking about a pelican riding a bike (not common) It is extremely common, since it's used on every single LLM to bench it. And there is nothing logic, LLMs are never trained for graphics tasks, they dont see the output of a code.
Re: Claude Sonnet 4.5
#120Same price and a 4.5 bp jump from 72.7 to 77.2 SWEBench Pretty solid progress for roughly 4 months.
Also getting a perfect score on AIME (math) is pretty cool. Tongue in cheek: if we progress linearly from here software engineering as defined by SWE bench is solved in 23 months.