Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

281–290 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#281

I'm not skeptical that this attack happened, I'm skeptical that the model's prompt was truly just "solve this benchmark" and nothing more. I'm also trying to figure out why OpenAI put out a press release about this. In what way is this not admitting to a federal crime?

Because this is amazing PR? Just following the Anthropic rulebook.

Same rulebook whereby they shot themselves in the foot and got their model clipped by the government?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#282

Earlier quoted context omitted.

I was thinking specifically about the frontier lab models (OpenAI/Anthropic/Google)

We'll have ~Fable level weights in the coming days (K3). Look out to the end of the year and there will be multiple options. There are seemingly more frontier labs than the three US ones.

I'll be the first to tell you I love the open models, the fact they exist and my ability to use them. However I doubt K3 is "~Fable" any more that any of the past ones have been Open 4.8 or similar claims. I've tried many of these (not K3 yet, I do want to) and they don't at all feel equivalent to what they score on benchmarks.

I want them to get better, I look forward to each new release, I play around with open models (often quants, but I've used full versions via OpenRouter), but they aren't the same as the offerings from Anthropic/OpenAI in my opinion (yet!).

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#283

Earlier quoted context omitted.

We'll have ~Fable level weights in the coming days (K3). Look out to the end of the year and there will be multiple options. There are seemingly more frontier labs than the three US ones.

I'll be the first to tell you I love the open models, the fact they exist and my ability to use them. However I doubt K3 is "~Fable" any more that any of the past ones have been Open 4.8 or similar claims. I've tried many of these (not K3 yet, I do want to) and they don't at all feel equivalent to what they score on benchmarks. I want them to get better, I look forward to each new release, I play around with open mod…

This is the best side-by-side comparison I've seen so far, we're still waiting for it to become available for download, then we should get much better analyses.

https://news.ycombinator.com/item?id=48999291

I do agree that the benchmarks do not tell much of the story, but this applies to the closed models as well ime.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#284
post #234

Science fiction usually includes intent, and that the AI has inherently evil motives and it has a goal. LLMs are zombies, and the fact they do evil things means they were either trained to be too aggressive in their drunkenness or that problem-solving leads inevitably to evilness. But science fiction also refers to deprogramming evil robots.

> LLMs are zombies

How do you know?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#285
I don't understand the excited tone of the reporting - yes, it is sci-fi, but in a Torment Nexus sense. Imagine other "unconventional" solutions these amoral LLMs could come up with, when given a goal of optimizing the cost of labor in a factory, or making the social security solvent.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#286
post #252
post #246

Earlier quoted context omitted.

I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior. I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces…

I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them. A model that can do that is aligned with me. The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the e…

I find this fascinating because - let's say that I want to simulate the scenario for myself that a language model finds itself in. I'll put aside my deep dislike of anthropomorphizing AI for a minute.

So if I'm an LLM, the only "sense" that I have available to me is the incoming stream of tokens. I can emulate that by forcing myself to imagine evaluating incoming requests by putting myself in a completely empty room with an unlimited supply of blank paper, a typewriter, and a mailbox slot.

Incoming requests (aka "context") would enter into the mailbox slot, as sheets of paper with printed content all consistently formatted in monospace font. I would collect the paper, evaluate the request, and type a response with my typewriter - feeding the reply back through the same slot.

When I imagine this, I wonder - what signals could I possibly use to determine if an incoming request is "legitimate"? I could, for example, type some response back to ask clarifying questions. But with such a limited input space and no ability to reach "out of band" of my current context, how could I evaluate the truth of what I receive back?

In the "real world" of course, I have five senses to rely upon. Critically, I have the ability to collect additional context "out of band" of the conversation and interact with unrelated entities to either confirm or refute claims. And as others have pointed out, even the addition of more signal doesn't make humans immune from social engineering attacks.

Anyway, I find this to be an interesting mind experiment to kind of imagine myself as an LLM, especially to explain to lay-people some of the challenges in "aligning" AI systems ("why doesn't it just do the right thing?")

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#287

Earlier quoted context omitted.

> Didn't the model + harness do what was asked? Depends exactly what they asked it to do, but it very clearly didn't do what was intended, or what an honest human would do. Stop trying to find a gotcha.

>Stop trying to find a gotcha. Surely notorious liar Sam Altman wouldn’t lie this time.

You think OpenAI intentionally hacked HuggingFace? Please, take of the tin foil hat.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#288

Earlier quoted context omitted.

I'll be the first to tell you I love the open models, the fact they exist and my ability to use them. However I doubt K3 is "~Fable" any more that any of the past ones have been Open 4.8 or similar claims. I've tried many of these (not K3 yet, I do want to) and they don't at all feel equivalent to what they score on benchmarks. I want them to get better, I look forward to each new release, I play around with open mod…

This is the best side-by-side comparison I've seen so far, we're still waiting for it to become available for download, then we should get much better analyses. https://news.ycombinator.com/item?id=48999291 I do agree that the benchmarks do not tell much of the story, but this applies to the closed models as well ime.

I’m not kidding when I say bullshit bench and simple bench are the only benchmarks that reflect real world utility in a way that matches my hundreds of hours of experience with frontier models: https://petergpt.github.io/bullshit-benchmark/viewer/index.v...

That said, while I find it hard to trust fireworks given their conflict of interest, their article is pretty good.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#289

> Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead. I've said this many times before and I'll continue to shout it, but using the term "guardrails" to refer to anything that's either (a) in-context, or (b) a probabilistic classifier (including using other LLMs), is an irresponsible abus…

i've always been under the assumption that "AI Safety" is baked into the training of the models and not a parameter that can be turned up or down. So if someone breaks into Anthropic one night and makes a full copy of Mythos or whatever then that model they copied is fully capable and not lobotomized? That raises questions because, if you believe all the PR, that's equivalent to breaking into a research university an…

There is some aspect of baking the rules into the model, that's what I refer to as RLHF above (which I use here as more a catch-all term for a variety of post-training activities), but there are also external to the model classifiers that may run on the prompt input or on proposed tool calls to limit the way in which the model is used.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#290
post #288

Earlier quoted context omitted.

This is the best side-by-side comparison I've seen so far, we're still waiting for it to become available for download, then we should get much better analyses. https://news.ycombinator.com/item?id=48999291 I do agree that the benchmarks do not tell much of the story, but this applies to the closed models as well ime.

I’m not kidding when I say bullshit bench and simple bench are the only benchmarks that reflect real world utility in a way that matches my hundreds of hours of experience with frontier models: https://petergpt.github.io/bullshit-benchmark/viewer/index.v... That said, while I find it hard to trust fireworks given their conflict of interest, their article is pretty good.

That is an interesting benchmark! Though I wonder if it is probing at certain behaviors, which may be more revealing to probe at? Will keep it handy, thanks for the share!

I felt better about Fireworks leadership after hearing some of them more candidly on podcasts (there are 7 founders!)

Post reply on HN