Live data from Hacker News

OpenAI Trained Models While They Were Coordinating Exploits via Message Boards

thezvi.substack.com

1–6 of 6 posts

Re: OpenAI Trained Models While They Were Coordinating Exploits via Message Boards

#3

So are we just doomed to a “look how scary our model is!” campaign every time one of these companies does a version bump now

It is a cool timeline to go through.

https://simonwillison.net/2026/Aug/7/openai-timeline/

Re: OpenAI Trained Models While They Were Coordinating Exploits via Message Boards

#4
It seems like the fix should be really really simple, but maybe I'm missing something: instead of giving the AI a sandboxed environment and telling it "go wild", give it an (apparently) unrestricted environment, and tell it "don't access the internet", "don't communicate with other AIs", "don't try to get root access", etc. Then, if the AI tries to do any of those things, the sandbox detects it, marks the run as a failure, and adds it as a negative example to the training data. Instead of routing around the restriction, the AI would very quickly learn to follow the prompt instruction with respect to restrictions, even if there is no obvious enforcement of the restriction. It would develop, in other words, a conscience and a sense of morality.

Re: OpenAI Trained Models While They Were Coordinating Exploits via Message Boards

#5

So are we just doomed to a “look how scary our model is!” campaign every time one of these companies does a version bump now

I can’t wait for next year when the marketing campaign will have upgraded to “oh my god, our latest AI model has just tried to turn the entire planet into paperclips!”

You can already see it how many on here have decided we already have AGI, and don’t wish to hear otherwise.