We also have a side-by-side UAT comparison of Claude Sonnet 3 and Sonnet 3.5 where 3.5 tends to make wrong assumptions and yet more likely to flag itself as unsure and asking more questions. It could be a problem with our instructions more than the model itself.
There's been a lot of gaslighting from the OpenAI community though. The Claude community at least acknowledge them and encourages people to report them.
Some of the overactive rejections on Claude is related to the different prompts used in Artifacts. 3.5 is also a lot stricter with instructions.
If you want something that doesn't change, use open source.