Seems to work again. But how does this happen in the first place. How could someone possibly have thought "hey I have an idea, let's put in a list of english words and just silently stop working if we have see even one of them in a substring". And people in this meeting would nod and say "yeah that sounds like an easy safety fix, let's do that". This just feels odd. This isn't a piece of forum software written by a 1…
They are not being randomly paranoid. Even if they did not have this fear, they would have rapidly developed it. We've all read the articles by muckraking journalists that take something an AI said and basically deliberately writes clickbait about how stupid or evil or worthless or whatever the AI is, even if the journalist had to filter through hundreds of replies (or, implicitly, by waiting for the dumbest stuff to rise to the top of social media, thousands or millions of replies) to get it. We've also read the articles where in someone uses the "fancy autocompleter", feeds it the moral equivalent of "Hey, how do you think you AIs will be taking over the world in five years?" and then is shocked, shocked at the "fancy autocompleter" filling in the yarn they are clearly asking for, and go running to either the media, or in particularly pathological cases, the academic literature making wild claims.
(I do not believe that "fancy autocompleter" is a complete description of LLMs, but in this particular case, it isn't a completely inaccurate mental model either. It shouldn't be a surprise that when you prompt it with X, you get more X.)
As a result the AIs are very heavily tuned to some combination of the political beliefs of the company writing them and the political beliefs dominant in the media coverage they are worried about, so they won't get very negative stories written about them. For this purpose, I'm taking the broadest possible definition of "political", not just "American politics in 202x", but the full range of "beliefs that not everyone agrees on and are things people are willing to exert some degree of power over". The AI companies have to take a stand, because taking a stand at least means someone can be on their side... if they just let the chips fall where they may they'll anger everyone because everyone can get the AI to say things that they in particular disagree with and they'll find themselves without friends. Unsurprisingly, the AI companies have been aligning their models with what they perceived to be the largest, most powerful political beliefs in their vicinity.
To be honest when I read them talking about "AI safety" I know they want me to be thinking "ensuring the AI doesn't take over the world or tell people to commit self harm" but what I see is them spending a lot of effort to politically align their AIs, with all that entails.