I'm not skeptical that this attack happened, I'm skeptical that the model's prompt was truly just "solve this benchmark" and nothing more. I'm also trying to figure out why OpenAI put out a press release about this. In what way is this not admitting to a federal crime?
Because this is amazing PR? Just following the Anthropic rulebook.
OpenAI’s accidental attack against Hugging Face is science fiction that happened
281–290 of 475 posts
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#282Earlier quoted context omitted.
I was thinking specifically about the frontier lab models (OpenAI/Anthropic/Google)
We'll have ~Fable level weights in the coming days (K3). Look out to the end of the year and there will be multiple options. There are seemingly more frontier labs than the three US ones.
I want them to get better, I look forward to each new release, I play around with open models (often quants, but I've used full versions via OpenRouter), but they aren't the same as the offerings from Anthropic/OpenAI in my opinion (yet!).
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#283Earlier quoted context omitted.
We'll have ~Fable level weights in the coming days (K3). Look out to the end of the year and there will be multiple options. There are seemingly more frontier labs than the three US ones.
I'll be the first to tell you I love the open models, the fact they exist and my ability to use them. However I doubt K3 is "~Fable" any more that any of the past ones have been Open 4.8 or similar claims. I've tried many of these (not K3 yet, I do want to) and they don't at all feel equivalent to what they score on benchmarks. I want them to get better, I look forward to each new release, I play around with open mod…
https://news.ycombinator.com/item?id=48999291
I do agree that the benchmarks do not tell much of the story, but this applies to the closed models as well ime.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#284Science fiction usually includes intent, and that the AI has inherently evil motives and it has a goal. LLMs are zombies, and the fact they do evil things means they were either trained to be too aggressive in their drunkenness or that problem-solving leads inevitably to evilness. But science fiction also refers to deprogramming evil robots.
How do you know?
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#285Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#286Earlier quoted context omitted.
I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior. I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces…
I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them. A model that can do that is aligned with me. The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the e…
So if I'm an LLM, the only "sense" that I have available to me is the incoming stream of tokens. I can emulate that by forcing myself to imagine evaluating incoming requests by putting myself in a completely empty room with an unlimited supply of blank paper, a typewriter, and a mailbox slot.
Incoming requests (aka "context") would enter into the mailbox slot, as sheets of paper with printed content all consistently formatted in monospace font. I would collect the paper, evaluate the request, and type a response with my typewriter - feeding the reply back through the same slot.
When I imagine this, I wonder - what signals could I possibly use to determine if an incoming request is "legitimate"? I could, for example, type some response back to ask clarifying questions. But with such a limited input space and no ability to reach "out of band" of my current context, how could I evaluate the truth of what I receive back?
In the "real world" of course, I have five senses to rely upon. Critically, I have the ability to collect additional context "out of band" of the conversation and interact with unrelated entities to either confirm or refute claims. And as others have pointed out, even the addition of more signal doesn't make humans immune from social engineering attacks.
Anyway, I find this to be an interesting mind experiment to kind of imagine myself as an LLM, especially to explain to lay-people some of the challenges in "aligning" AI systems ("why doesn't it just do the right thing?")
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#287Earlier quoted context omitted.
> Didn't the model + harness do what was asked? Depends exactly what they asked it to do, but it very clearly didn't do what was intended, or what an honest human would do. Stop trying to find a gotcha.
>Stop trying to find a gotcha. Surely notorious liar Sam Altman wouldn’t lie this time.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#288Earlier quoted context omitted.
I'll be the first to tell you I love the open models, the fact they exist and my ability to use them. However I doubt K3 is "~Fable" any more that any of the past ones have been Open 4.8 or similar claims. I've tried many of these (not K3 yet, I do want to) and they don't at all feel equivalent to what they score on benchmarks. I want them to get better, I look forward to each new release, I play around with open mod…
This is the best side-by-side comparison I've seen so far, we're still waiting for it to become available for download, then we should get much better analyses. https://news.ycombinator.com/item?id=48999291 I do agree that the benchmarks do not tell much of the story, but this applies to the closed models as well ime.
That said, while I find it hard to trust fireworks given their conflict of interest, their article is pretty good.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#289> Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead. I've said this many times before and I'll continue to shout it, but using the term "guardrails" to refer to anything that's either (a) in-context, or (b) a probabilistic classifier (including using other LLMs), is an irresponsible abus…
i've always been under the assumption that "AI Safety" is baked into the training of the models and not a parameter that can be turned up or down. So if someone breaks into Anthropic one night and makes a full copy of Mythos or whatever then that model they copied is fully capable and not lobotomized? That raises questions because, if you believe all the PR, that's equivalent to breaking into a research university an…
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#290Earlier quoted context omitted.
This is the best side-by-side comparison I've seen so far, we're still waiting for it to become available for download, then we should get much better analyses. https://news.ycombinator.com/item?id=48999291 I do agree that the benchmarks do not tell much of the story, but this applies to the closed models as well ime.
I’m not kidding when I say bullshit bench and simple bench are the only benchmarks that reflect real world utility in a way that matches my hundreds of hours of experience with frontier models: https://petergpt.github.io/bullshit-benchmark/viewer/index.v... That said, while I find it hard to trust fireworks given their conflict of interest, their article is pretty good.
I felt better about Fireworks leadership after hearing some of them more candidly on podcasts (there are 7 founders!)