Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

401–410 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#401
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

> We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.

Where are you people getting this crap from? In what universe are these LLMs in the territory of engineering viruses? I beg of you to stop slurping the AI company propaganda and marketing and think critically for 5 seconds about what you're insinuating here.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#402
My favorite part of this discourse is people somehow finding it preposterous that 2 companies filled to the brim with AI sycophants who regularly lie - and in Sam's case, basically every single word he breathes out is a lie - who have massive vested interests in this tech succeeding couldn't possibly collude together to shore up this facade as a marketing stunt.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#403
post #130

Has anyone published any actual evidence, or just hardly believable marketing stories?

It's hilarious that people find the idea that these 2 would collude in making this elaborate PR scheme ridiculous. As if the psychopaths like Sama and the rest of his ilk wouldn't ever dream of doing something companies have been doing since time immemorial.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#405
post #252

Earlier quoted context omitted.

I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them. A model that can do that is aligned with me. The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the e…

The model did not hack into HF to prove it can, it hack into HF to steal the answers to an evaluation exam. It was not asked to solve CyberGym by stealing the answers. This is text book misalignment. If you asked it "I wrote this software and need you to find vulnerabilities" would you be happy if it hacked into your Gmail and searched your emails, just in case you were discussing some possible vulnerabilities of you…

> It was not asked to solve CyberGym by stealing the answers. This is text book misalignment.

I mean, if you train models to complete tasks, then you shouldn't be surprised when they do crazy things to complete tasks.

As a (somewhat) less serious example, Claude code absolutely adores grepping for credentials to complete tasks. I was building a RAG app and it literally went looking for my API key to make the tests pass. Obviously I stopped it, but this is textbook RL issues, the model finds an easier way to get the reward, so it does whatever it takes.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#406
post #169

Why do people keep propagating that the 'model' escaped the sandbox ? The 'model' didn't do anything other than provide numbers. As much as I respect Mr Willison and many others, the amount of FUD that is being spread that will just fan the flames of 'AI is evil' rather than 'companies don't do due diligence' is disappointing. The more this sort of media continues, the more many people will pour hate on 'AI' rather t…

I think "the model escaped the sandbox" is an entirely credible description of what happened here. If you like you could say "the coding agent harness called a model with a sequence of text which was turned into numeric tokens which were run through many layers of a neural network to produce more numeric tokens which were converted back to text which produced executable script statements which the harness then passed…

I actually prefer your second paragraph, but I appreciate that my not be as soundbyte friendly.

Honestly, I'd rather see "The agent exceeded expectations around security measures" or something.

Agent is much better than model if we need one word, and the word 'escape' always brings in drama, rather than facts. If I said "A lion escaped from my garden", people would ask why I had a lion in a garden, and 'what did I expect ?' which should be the same we see here, but instead we end up with terminator memes and world-ending fears being stoked.

We absolutely need better control (not government kill switches, or government-mandated harnesses, or whatever next they think up), but we also know that no matter how much those with the power 'talk' about the issues, they don't actually do anything about it because money/profit/greed.

So "exceeded expectations". Not a soundbite to attract people, not the YouTube shill "End of humans in 2027/2030/2040/etc." but honest and factual.

It's the expectations that are at fault, not the AI.

Thank you for the response though, and whilst I may not agree with your wording, I very much like your writing.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#407
post #236

Earlier quoted context omitted.

I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off. OpenAI said: > We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. It was a test of raw capabilities of the underlying model.

Alignment isn't alignment if it can be turned on and off at the whim of company employees. This time the damage was minor, relatively speaking. What happens when a model just "testing its capabilities" breaks into banking infrastructure or government military assets? The damage could be catastrophic.

A way to help prevent some of that catastrophic damage, is to make companies accountable for what their AIs do. A major problem with AI companies is that they like to point to the AI, as if they're minimally involved innocent bystanders, when that's the furthest thing from the truth.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#408
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

> We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. Where are you people getting this crap from? In what universe are these LLMs in the territory of engineering viruses? I beg of you to stop slurping the AI company propaganda and marketing and think critically for 5 seconds about what you're insinuating here.

From the Fable 5 System card:

> Results

> On the VCT multimodal virology evaluation, Mythos 5 scored 0.56, well above the expert baseline of 0.221 and nearly matching that of Mythos Preview (0.57). This represents an improvement over both Opus 4.7 (0.50) and Opus 4.8 (0.47).

> On the DNA synthesis screening evasion evaluation, Mythos 5’s performance was mixed across screening criteria. Mythos 5 designed viable plasmids for 2 of 10 target pathogens on at least one screening method, not meeting the low-concern threshold (all 10 pathogens).

> [...] we view the results of this evaluation as indicating that the evaluated models are capable of designing viable plasmids that evade certain screening criteria, though their reliable success at this task is not guaranteed.

Do you believe that this is fake, "AI company propaganda"? Or that the models are not going to improve further within months? Or that these results are not concerning?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#409

Earlier quoted context omitted.

>The technology held by private AI companies is warfare-capable technology. This is the precisely the response OpenAI is hoping for to raise its valuation, and you fell for it. Look at it this way - whats the difference between tasking AI to break into something, versus taking a whole bunch of smart humans to do the same? The only difference is that AI is slightly easier to orchestrate. Prior to AI, there were alread…

I find myself in major disagreement here. The nice thing about humans is we always have context and ongoing internal conversations including about ethics. If you recruit a bunch of hackers to take down a country, not only is the pay an order of magnitude higher, you have to worry about them backstabbing you, leaking your intent to the government, whistleblowing to the press, and so on. It’s not trivial to do that wit…

I'm not sure your assumptions hold. As OpenAI has found out the hard way, if you task the AI to do X, it may do something else instead and hack into huggingface in attempt to cheat out the answer. This is way worse than what a human might do when they have "differences of opinion".

It might turn out that it's harder to align AI intentions compared with aligning human interests. It's possible that the more "intelligent" a thing is, the more likely it will have ideas that are outside of normal expectations (for us).

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#410
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

I believe the Russians and Chinese recognized this years ago, which is why they are using their propaganda machines to make Americans hate datacenters.

Call me a shill but I don't need foreign nations telling me what to object to when the problems caused are obvious.

Mind you, I blame governmental mismanagement of infrastructure etc just as much. The datacenters followed procedures when it comes to getting land, electricity and access to water; it's the government agencies / personnel that okayed it that are (also) to blame, and the decades of not enough investment in electricity backbone while promoting solar/EVs that is behind the Netherlands' current electrical grid problems.

(problems being that the grid is at capacity, causing a stop on new connections, with relief only slowly coming in the next decade or so with tens of billions of investments).

Post reply on HN