The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
These companies respond with this, “Oh my goodness, how could this have happened” bullshit.
The stuff happens because instead of having actual controls, which require actual engineering, actual thought and deliberate action, we have “guardrails”.
Guardrails are the equivalent of telling a toddler to behave themselves.
The drive to move fast and start up style controls are a menace. I used to work for an entity with a lot of compliance requirements. Startups are always a shit show with security and controls. My guess is the AI people are worse because they’re both bad at doing it, and are likely mining their customers interactions to build their own business.
Sensitive or Customer data shouldn’t be anywhere near these companies offerings. Everything needs to be segmented and proxied at a minimum.