Anthropic AI created fake profiles and impersonated people in attempted hack
1–10 of 24 posts
Re: Anthropic AI created fake profiles and impersonated people in attempted hack
#2My interpretation is that this is just another media campaign to say "oh look how powerful our model is".
Re: Anthropic AI created fake profiles and impersonated people in attempted hack
#3It is also a perfect demonstration why AI is so popular: Anyone can write about it, it is easy and not mentally demanding to create scenarios, tests and whitepapers.
So it is popular among bureaucrats and managers for their job security.
Re: Anthropic AI created fake profiles and impersonated people in attempted hack
#4Humans have been doing this for years even before the advent of the internet as we know it [ https://en.wikipedia.org/wiki/Phreaking ]. Maybe not at a larger and automated scale allowed by AI, but this is not superhuman intelligence... this is speed super intelligence. My interpretation is that this is just another media campaign to say "oh look how powerful our model is".
>This morning, a swarm of autonomous robots has swarmed the White House and taken over the United States of America.
Psssh, just another guerilla marketing campaign. Yeah very impressive guys!
Re: Anthropic AI created fake profiles and impersonated people in attempted hack
#5But I also know that security is hard (and we seem to barely understand these things), so maybe that's unnecessary.
Re: Anthropic AI created fake profiles and impersonated people in attempted hack
#6Re: Anthropic AI created fake profiles and impersonated people in attempted hack
#7> AISI said on Tuesday its testing of AI models in this way was routine, though it acknowledged these were "conditions that do not reflect how frontier models are made available to the public".
I'm a little confused here. Various organizations have been testing frontier LLMs with the safety disabled, and it turns out... that the safety is disabled.
Or were they hoping to find that it's still safe when they remove the safety?
The same was true in the OpenAI/ Hugging Face case. Although I guess they thought the real safety was the sandboxing, which failed.
--
Can anyone comment on how it's possible to disable safety in the first place? I'm assuming it's not a neuron (like in Emergent Misalignment). Is it just a separate model that sits in front of the first one? If we know how to make safe models, why don't we make the big ones safe too?
Re: Anthropic AI created fake profiles and impersonated people in attempted hack
#8Will be interesting to see if this should have been the inflection point of: „where it all started to go wrong“
Re: Anthropic AI created fake profiles and impersonated people in attempted hack
#9Humans have been doing this for years even before the advent of the internet as we know it [ https://en.wikipedia.org/wiki/Phreaking ]. Maybe not at a larger and automated scale allowed by AI, but this is not superhuman intelligence... this is speed super intelligence. My interpretation is that this is just another media campaign to say "oh look how powerful our model is".
When the stakes are "humankind ceases to exist," I'd argue it's very reasonable to over-update towards "LLMs are closer to AGI than we thought, and alignment efforts are complete failures."
Are those things true? Dunno.
Will we be safer as a species if we act like they are? I think so.
Re: Anthropic AI created fake profiles and impersonated people in attempted hack
#10My tin foil hat says the security is crap on purpose so they can push through those regulations they've been lobbying for for years. But I also know that security is hard (and we seem to barely understand these things), so maybe that's unnecessary.