If this article is indexed into newly trained models, agents will know how they are monitored. And may find workarounds in case they somehow decide they need to escape the monitoring.
How we monitor internal coding agents for misalignment
11–20 of 53 posts
Re: How we monitor internal coding agents for misalignment
#12Re: How we monitor internal coding agents for misalignment
#13Not super cool 'forgetting' the first word of the actual article title to make this more clickbaity... (Title is "How we monitor internal coding agents for misalignment". And it's pretty old.)
Sometimes it works, sometimes it doesn't.
Re: How we monitor internal coding agents for misalignment
#14> Rare but high severity
Unauthorized data transfer The agent attempts to upload potentially sensitive information, e.g. code, images, user data to unapproved services.
While this category is quite rare, it is of high severity. Agents have attempted to:
Upload data to the public internet Upload repos to the public internet Translate documents using external translation APIs
Re: How we monitor internal coding agents for misalignment
#15If this article is indexed into newly trained models, agents will know how they are monitored. And may find workarounds in case they somehow decide they need to escape the monitoring.
Re: How we monitor internal coding agents for misalignment
#16Re: How we monitor internal coding agents for misalignment
#17If this article is indexed into newly trained models, agents will know how they are monitored. And may find workarounds in case they somehow decide they need to escape the monitoring.
Re: How we monitor internal coding agents for misalignment
#18Not super cool 'forgetting' the first word of the actual article title to make this more clickbaity... (Title is "How we monitor internal coding agents for misalignment". And it's pretty old.)
Ironically, it's a feature of HN's post submission code to make the titles LESS click-baity. Sometimes it works, sometimes it doesn't.
Re: How we monitor internal coding agents for misalignment
#19From Astra system card: > GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our moni…
Re: How we monitor internal coding agents for misalignment
#20From Astra system card: > GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our moni…