Earlier quoted context omitted.
Even if this were all true, it points to a fundamental risk of using LLM's for important tasks, which is that it is not at all clear to a user that this prompt would cause a problem. The LLM doesn't say "I'm sorry Dave, I just can't do that", it just complies with it and gets the wrong answer. You can always make excuses for the LLM afterwards, but software with hidden risks like this would not be considered good or…
People really need to stop trying to model an LLM as some kind of magical software component: it all makes a lot more sense if you model it as an under-performing poorly-aligned employee; so like, maybe a distracted kid working for peanuts at your store. You wouldn't trust them to with all of your money and you wouldn't trust them to do a lot of math--if they had to be in charge of checkout, you'd make sure they are…
I once gave a 10-dollar bill to a young man serving at the cashier at a store, and he gave me 14 dollars back as a change. I pointed out that this made no sense. He bent down, looked closer at the screen of his machine, and said "Nope, 14 dollars, no mistake". I asked him if he thought I gave him 20. He said no, and even shown me the 10-dollar bill I just gave him. At that point I just gave up and took the money.
Now that I think about it, there was an eerie similarity between this conversation and some of the dialogues I had with LLMs...