Unlike most other commenters, I applaud him for acting on his principles. If you sincerely believe that, of course you should act. You might not succeed, but your voice might be the one that tips the scales and starts a broader movement. This doesn't mean I agree with him. The fears of doomsday caused by rapid takeoff have been with us since day 1 and the mechanism is always basically "AI invents magic that sets it f…
To borrow on the 1990s Slashdot meme: 1. Invent transformer architecture. 2. Scale it up. 3. ??? 4. Machines become sentient and kill us all. OpenAI and Anthropic pinky promise that they have figured out #3 and they're not BSing just to get more funding, no. But because we live in a culture of fear, everyone eats it up no questions asked.
At lower capability levels, the patterns are very clear and have been studied to death. E.g. Why LLMs say they have correctly fixed a broken test when they haven't. What we saw with Hugging Face is literally the exact same problem, just scaled up and with more capable agents. This shit was predicted decades ago...
No one can say exactly how it will play out as the complexity increases, but the risks are becoming extremely obvious.
I personally think it's extremely unlikely to "kill everyone", but there are many outcomes far short of that which seem quite plausible and rather undesirable. Russian roulette is not a smart game.
If you read the METR report and aren't scared at all, then I'd love to know why. It would help me sleep better. So please share.
TL;DR; increasing capabilities, reward hacking, and unsafe training regimes.