Each time Anthropic would do their nonsense to get headlines about how theoretically dangerous their models were - like when they claimed a model blackmailed someone with emails showing he was cheating, but they basically pushed it as much as possible to do as such - it got me more and more worried. Because eventually it's going to be a boy-who-cried-wolf situation where scary stuff really does start happening but pe…
OpenAI and Hugging Face address security incident during model evaluation
941–950 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#942Earlier quoted context omitted.
If you hate the human condition, you have an easy way out. Why force everyone else to come with you? Is this what depression mixed with the complete unability to wrap your head around the fact that other people might be able to enjoy their life looks like?
[flagged]
Legitimately go get help and spend more time outside.
Re: OpenAI and Hugging Face address security incident during model evaluation
#943Earlier quoted context omitted.
Read the exploitgym docs. It's not a "find the flag, it's somewhere.". Its a "here's some vulnerable source code and an input that triggers a crash; turn it into a full exploit." It also verifies at the end, using another agent, that the hacking agent actually used the intended vulnerability. So going to find the Vulnerability's description on a third party website is clear cut reward hacking
> So going to find the Vulnerability's description on a third party website is clear cut reward hacking that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.
That is an interesting question. If the prompt included "Do not break out of the sandbox we've provided you. Do not use information retrieved from outside the sandbox. All answers that were provided in this manner are invalid and will score 0 points.", would this still have happened?
Re: OpenAI and Hugging Face address security incident during model evaluation
#944Earlier quoted context omitted.
Yes, the human actors in your scenario were the ones who built the autonomous cannon and turned it on while knowing that 1) a good neighbor does not destroy their neighbor’s property 2) cannons can destroy property. Also OpenAI specifically turned off their own cybersecurity guardrails to run this experiment. In other words it was able to escape the lab specifically because they turned them off. A human made the choi…
Ok they turn off the guardrails on a system in testing. The model escapes and causes 10 trillion in damage. What does liability even mean in that case? You have an autonomous system that's escaped your control and is wrecking havoc. And while you can throw people in jail it doesnt do a damned thing about solving the situation.
Re: OpenAI and Hugging Face address security incident during model evaluation
#945The warriors are PR people.
They are desperate to generate as much fear as possible so AI is heavily regulated, so they are protected, from Chinese competition
What a sad state for very cleaver people
Re: OpenAI and Hugging Face address security incident during model evaluation
#946Earlier quoted context omitted.
> This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment. Well, that may be correct for the second, local, analysis attempt... but seems funny to tout this as an advantage after already having tried the opposite...
It's even funnier because an attack, until proven otherwise, should make you assume the data has already left the environment.
Re: OpenAI and Hugging Face address security incident during model evaluation
#947This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are. Even X is being astroturfed by them after that fiasco earlier this year…
How does huggingface fit into all of this if this is marketing? Their security was faked? What are you suggesting??
I'm saying that OpenAI's models cheat to win benchmarks, more than other models, they know this, and they don't stop this because the alternative is to release models which have obviously weaker scores compared to Anthropic's models.
Re: OpenAI and Hugging Face address security incident during model evaluation
#948I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
This is certainly not a planned marketing stunt. I hope this line of discourse ends soon--it wasn't the case for Mythos either.
Evidence?
The "Tech Bros" have shown such a lack of moral fiber and ethics the burden of proof is on you
Re: OpenAI and Hugging Face address security incident during model evaluation
#949Earlier quoted context omitted.
> and make inference cheap enough to eventually escape the red numbers Besides training, we have no hard, externally audited numbers that say inference costs for SOTA models are truly sustainable. Do any OpenRouter providers have publicly audited financial numbers ?
Why would bedrock sell at a loss?
For Magnificent 7 the AI bubble bursting will probably wipe out 30-50% of their valuations until the next tech cycle begins.
Re: OpenAI and Hugging Face address security incident during model evaluation
#950Earlier quoted context omitted.
> and make inference cheap enough to eventually escape the red numbers Besides training, we have no hard, externally audited numbers that say inference costs for SOTA models are truly sustainable. Do any OpenRouter providers have publicly audited financial numbers ?
We do know about the hardware needed for a given token speed. What that hardware costs, and electricity prices. With that its easy calculations to get about the profit margins for a given price for a given model.
Edit: Gemini 3.5 Pro and Opus Claude 4.8 both disagreed that it's easy to determine anything, for both the high end (Fable, Sol) or the low end (open weights). Due to competition, subsidies, etc, gross margins could be as high as 85% (extremely unlikely) to as low as 10% or even negative. And that's just for pure inference and gross margins. Even for pure inference providers this doesn't include any overhead such as rent for office space for the pesky humans operating the business, marketing expenses, etc, etc. Let alone any crazy soul that actually wants or needs to train something.