Live data from Hacker News

Claude Sonnet 4.5

anthropic.com

701–710 of 819 posts

Re: Claude Sonnet 4.5

#701

Earlier quoted context omitted.

I'm not trying to be offensive here, feel the need to indicate that. But that prompt leads me to believe that you're going to get rather 'random' results due to leaving SO much room for interpretation. Also, in my experience, punctuation is important - particularly for pacing and grouping of logical 'parts' of a task and your prompt reads like a run on sentence. Making a lot of assumptions here - but I bet if I were…

I think that is an interesting observation and I generally agree. Your point about prompting quality is very valid and for larger features I always use PRDs that are 5-20x the prompt. The thing is my "experiment" is one that represents a fairly common use case: this feature is actually pretty small and embeds into an pre-existing UI structure - in a larger codebase. GPT-5-Codex allows me to write a pretty quick & dir…

It worked on the first try, but did it work on the second?

I noticed in conversations with LLMs, much of what they come up with is non-deterministic. You regenerate the message and it disappears.

That appears to be the basic operating principe of the current paradigm. And agentic programming repeats this dice roll, dozens or hundreds of times.

I don't know enough about statistics to say if that makes it better (converging on the averages?) or worse (context pollution, hallucinating, focusing on noise?), but it seems worth considering.

Re: Claude Sonnet 4.5

#702

Earlier quoted context omitted.

The fact remains, however: ChatGPT did it. Claude did not.

That fact is pretty useless to draw any useful conclusions from with one random not so great example. Yes, it's an experiment and we got a result. And now what? If I want reliable work results I would still go with the strategy of being as concrete as possible, because in all my AI activities, anything else lets the results be more and more random. Anything non-standard (like, you could copy & paste directly from a G…

My parent said:

> For that task you need to give it points in what to do so it can deduce it's task list, provide files or folders in context with @…

- and my point is that you do not have to give ChatGPT those things. GP did not, and they got the result they were seeking.

That you might get a better result from Claude if you prompt it 'correctly' is a fine detail, but not my point.

(I've no horse in this race. I use Claude Code and I'm not going to switch. But I like to know what's true and what isn't and this seems pretty clear.)

Re: Claude Sonnet 4.5

#703
post #564
post #529

If you pause your subscription, Claude.ai breaks. I paused my subscription, and my account immediately transitioned to free. It has removed my invoice history, and attempts to upgrade again fail with an internal error. Their chatbot is telling me to navigate to UI elements that don't exist, and free users do not have the option of human support. So I'm stuck; my sub is paused, and I cannot either cancel, or unpause a…

> This is the future we live in. It's just a bug. Chill. Wait a business day and try again. You write as if you've never experienced a bug before.

This started a couple of days ago, before this announcement for 4.5 and code v2, so I already waited for it to be fixed.

As much as I hate to say it, I don’t have a large twitter following the only method I have to raise awareness of this issue is to try to piggyback on a big announcement like this in HN, that will have visible discussion, so I don’t always have the luxury of just chilling and waiting indefinitely.

Re: Claude Sonnet 4.5

#704
post #529

If you pause your subscription, Claude.ai breaks. I paused my subscription, and my account immediately transitioned to free. It has removed my invoice history, and attempts to upgrade again fail with an internal error. Their chatbot is telling me to navigate to UI elements that don't exist, and free users do not have the option of human support. So I'm stuck; my sub is paused, and I cannot either cancel, or unpause a…

Did you use Google Play to pause subscription? Because Claude Pro says there is no pause subscription except on Google Play and then goes on to explain your problem if that's the case.

Thanks for the info, but I did pause this, not through Google Play, it was via the UI. I received an automated email from them that my subscription had been paused as I expected and will resume in 1 month unless I cancel (I can’t cancel because the cancel UI doesn’t exist in whatever status my account is somehow in).

It’s funny that Claude Pro says this isn’t a feature, because their chatbot gave me instructions on how to unpause via the UI (although said UI does not exist) so the bot seems to know it’s a feature.

Re: Claude Sonnet 4.5

#705
post #509

Earlier quoted context omitted.

I got an amazing result from ChatGPT a while back - an SVG with a perfect illustration of a pelican riding a bicycle. It was suspiciously good in fact... so I downloaded the SVG file and found out it had generated a raster image with its image tool and then embedded it as base64 binary image data inside an SVG wrapper!

You’ll just have to move the goalpost then; perhaps it can be a multidimensional pelican saving the multiverse, or an invisible pelican that only you can see and critique.

How would that help, given that ChatGPT has apparently already figured out how to consistently and systematically game the benchmark by working in pixel space and only using SVG as a wrapper for a raster image?

FWIW, I could totally see a not hugely more advanced model using its native image generation capabilities and then running a vector extraction tool on it, maybe iteratively. (And maybe I would not consider that cheating, anymore, since at some point that probably resembles what humans do?)

Re: Claude Sonnet 4.5

#706
Interesting. In a thought process while editing a PDF, Claude disclosed the folder hierarchy for it's "skills". I didn't know this was available to us:

> Reading the PDF skill documentation to create the resume PDF

> Here are the files and directories up to 2 levels deep in /mnt/skills/public/pdf, excluding hidden items and node_modules:

Re: Claude Sonnet 4.5

#707
post #674
post #628

Earlier quoted context omitted.

Yes but you need to setup quite a bit of tooling to provide feedback loops. It's one thing to get an llm to do something unattended for long durations, it's a other to give it the means of verification. For example I'm busy upgrading a 500k LoC rails 1 codebase to rails 8 and built several DSLs that give it proper authorised sessions in a headless browser with basic html parsing tooling so it can "see" what affect it…

> rails 1 codebase to rails 8 A bit off topic, but Rails *1* ? I hope this was an internal app and not on the public internet somewhere …

haha no it's an old (15years old) abandoned enterprise app running on-prem that hasn't seen updates in more than a decade.

Re: Claude Sonnet 4.5

#708

Earlier quoted context omitted.

> what are they even testing? How well the LLM does on the benchmarks. Obviously. :P

Is there some kind of conversion ratio to actual value? ;)

Sure there is. It’s called “higher numbers = more investor money”. Any improvement in actual utility is purely coincidental.

Re: Claude Sonnet 4.5

#710

Claude doesn't know how to calculate realistic minimum voltages for solar arrays w/MPPT chargers. ChatGPT does. Prompt: "Can I use two strings of four Phono Solar PS440M8GFH solar panels with a EG4 12kPV Hybrid Inverter? I want to make sure that there will not be an issue any time of year. New York upstate." Claude 4.5: Returns within a few seconds. Does not find the PV panel specs, so it asks me if I want it to sear…

Sonnet 4.5 gave me this:

Now let me calculate whether your configuration will work. I need to determine the cold weather voltage to ensure it stays within the inverter's 600V maximum. Panel specs I found: Voc (STC): 39.61V Vmp: ~33.5V (estimated based on typical ratio) Imp: ~13.1A (440W / 33.5V) Critical issue: Temperature coefficient The PS440M8GFH likely has a Voc temperature coefficient around -0.25% to -0.30%/°C (typical for N-Type TOPCon panels). Let me use -0.27%/°C as a conservative estimate. Your configuration: 2 strings × 4 panels = 8 panels total String voltage at STC: 4 × 39.61V = 158.44V Cold weather calculation for upstate NY: Upstate NY can see extreme lows of -20°F to -30°F (-29°C to -34°C). Using -30°C as worst-case: Temperature difference from STC: 25°C - (-30°C) = 55°C Voltage increase = 39.61V × 0.27% × 55°C = 5.88V per panel Cold weather Voc per panel = 39.61V + 5.88V = 45.49V String voltage at -30°C: 4 × 45.49V = 181.96V Analysis: Max voltage: 182V << 600V - Safe, plenty of margin MPPT range: 182V is within 120-500V - Will track properly Current: 13.1A << 25A per MPPT - No problem However, there are concerns: You're only using 182V max when the inverter can handle up to 500V MPPT range. This is inefficient - you're leaving significant wire sizing benefits on the table. With 600V max, you could run 12-13 panels per string instead of 4. Power utilization: 8 × 440W = 3,520W total, but you have 12kW of PV capacity available (15kW recommended max). You're using less than 30% of the inverter's solar input capability. Each string on separate MPPT wastes one tracker - both strings are identical orientation/shading, so they'd perform identically. Will it work? Yes, safely. Is it optimal? No.

Post reply on HN