Live data from Hacker News

Claude Sonnet 4.5

anthropic.com

691–700 of 819 posts

Re: Claude Sonnet 4.5

#691
post #529

If you pause your subscription, Claude.ai breaks. I paused my subscription, and my account immediately transitioned to free. It has removed my invoice history, and attempts to upgrade again fail with an internal error. Their chatbot is telling me to navigate to UI elements that don't exist, and free users do not have the option of human support. So I'm stuck; my sub is paused, and I cannot either cancel, or unpause a…

I feel like we're just renting our digital lives.

You are.

It’s the same reason why many are becoming evangelists of hosting their own email, note apps, etc.

Re: Claude Sonnet 4.5

#692

Earlier quoted context omitted.

GPT-5 is like the guy on the baseball team that's really good at hitting home runs but can't do basic shit in the outfield. It also consistently gets into drama with the other agents e.g. the other day when I told it we were switching to claude code for executing changes, after badmouthing claude's entirely reasonable and measured analysis it went ahead and decided to `git reset --hard` even after I twice pushed back…

Gemini is an excellent collaborator? It’s the one AI that keeps telling me I’m wrong and refuses to do what I ask it to do, then tells me “as we have already established, doing X is pointless. Let’s stop wasting time and continue with the other tasks” It’s by far the most toxic and gaslighting LLM

What you get when you mix Google's excellent technical background with interoffice politics and extreme political correctness

Re: Claude Sonnet 4.5

#693

Anecdotal evidence. I have a fairly large web application with ~200k LoC. Gave the same prompt to Sonnet 4.5 (Claude Code) and GPT-5-Codex (Codex CLI). "implement a fuzzy search for conversations and reports either when selecting "Go to Conversation" or "Go to Report" and typing the title or when the user types in the title in the main input field, and none of the standard elements match, a search starts with a 2s de…

Interesting, in my experience Claude usually does okay with the first pass, often gets the best visual/ui output, but cannot improve beyond that even with repeated prompts and is terrible at optimising, GPT almost the opposite.

Re: Claude Sonnet 4.5

#694
It took me one question to have it spit out a completely dreamt up codebase, complete with emojis, promises of solutions and fixing all my problems, and of course nothing of it worked. It was a very simple question about something very well documented (Oban timeouts).

I doubt LLM benchmarks more and more, what are they even testing?

Re: Claude Sonnet 4.5

#695

It took me one question to have it spit out a completely dreamt up codebase, complete with emojis, promises of solutions and fixing all my problems, and of course nothing of it worked. It was a very simple question about something very well documented (Oban timeouts). I doubt LLM benchmarks more and more, what are they even testing?

> what are they even testing?

How well the LLM does on the benchmarks. Obviously.

:P

Re: Claude Sonnet 4.5

#696
post #488
post #377

Earlier quoted context omitted.

They need to benchmaxxx a whole lot harder, the illustrations still all universally suck!

I fully expect a model to output a SVG made up of 1000x1000 rectangles (i.e. pixels) representing a raster image of a beautifully hand-drawn pelican riding a bicycle any day now :)

ive got such pixelated rectangle SVG's a few times.

also with cursor, "write me a script that outputs X as an svg" it has given me rectangles a few times.

Re: Claude Sonnet 4.5

#697
post #143

Earlier quoted context omitted.

> it went ahead and decided to `git reset --hard` even after I twice pushed back on that idea So this is something I've noticed with GPT (Codex). It really loves to use git. If you have it do something and then later change your mind and ask it to undo the changes it just made, there's a decent chance it's going to revert to the previous git commit, regardless of whether that includes reverting whole chunks of code i…

Just to add another anecdotal data point, ive absolutely observed Claude Code doing exactly this as well with git operations.

I've gotten the `git reset --hard` with Claude Code as well, just not immediately after (1)) explicitly pushing back against the idea or (2) it talking a bunch of shit about another agent's totally reasonable analysis.

Re: Claude Sonnet 4.5

#698

It took me one question to have it spit out a completely dreamt up codebase, complete with emojis, promises of solutions and fixing all my problems, and of course nothing of it worked. It was a very simple question about something very well documented (Oban timeouts). I doubt LLM benchmarks more and more, what are they even testing?

> what are they even testing? How well the LLM does on the benchmarks. Obviously. :P

Is there some kind of conversion ratio to actual value? ;)

Re: Claude Sonnet 4.5

#699

Earlier quoted context omitted.

I'm thinking about switching to ChatGPT Pro also. Any idea what maxes it out before I need to pay via the API instead? For context I'm using about 1b tokens a month so likely similar to you by the sounds of things.

On pro tier have not been able to trigger the usage cap. Pro Local tasks: Average users can send 300-1,500 messages every 5 hours with a weekly limit. Cloud tasks: Generous limits for a limited time. Best for: Developers looking to power their full workday across multiple projects.

Thank you, that's very helpful. I think I could get close to that in some coding sessions where I'm running multiple in parallel but I suspect it's very very rare. Even with token efficient gpt5-codex my OpenAI bill is quite high so I think I will switch to Pro now.

Re: Claude Sonnet 4.5

#700

I need to try Claude - haven't gotten to it. I use AI for different things, though, including proofreading posts on political topics. I have run into situations where ChatGPT just freezes and refuses. Example: discussing the recent rape case involving a 12-year-old in Austria. I assume its guardrails detect "sex + kid" and give a hard "no" regardless of the actual context or content. That is unacceptable. That's like…

I'd imagine that the proportion of "legit" conversations around these topics and those that they're intending to not allow is large enough that it doesn't make sense for them to even entertain the idea of supporting those conversations. As a rather hilarious and really annoying related issue - I have a real use where the application I'm working on is partially monitoring/analyzing the bloodlines of some rather specif…

This is the result of Anthropic and others focusing on imaginary threats about things the model cannot realistically do - such as engineer bio weapons.

To guard against the imaginary threats, they compromise real use cases.

Post reply on HN