Earlier quoted context omitted.
This was clearly explained by OpenAI in their initial press release on 7/21 [0]: > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. […] The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test s…
It does not explain how agents months later would "collaborate" to hack Hugging Face.
Timeline of the OpenAI accidental attack against Hugging Face
181–190 of 287 posts
Re: Timeline of the OpenAI accidental attack against Hugging Face
#182Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…
Because that is fundamentally impossible given how they work...
The thing does not even know when it succeeds or fails. Actually the thing does not "know" at all...
All it can does is to show some limited textual behavior that matches with "knowing"..
Re: Timeline of the OpenAI accidental attack against Hugging Face
#183Earlier quoted context omitted.
Beyond a whole lot of online conspiracy theories I haven't seen anything that suggests to me that OpenAI aren't not telling the truth about what happened here. I find the Black Hat presentation in particular very credible. Also the Hugging Face technical report. (As an example of something I don't find credible: https://openai.com/index/responding-next-frontier-critical-c... is a total nothing burger. It's the other…
[edit] I've now watched the video on the idea that your write-up was misleading. BUT the video is much worse. For two months with highly dangerous agents agents were hacking a service and none of the researchers watched (drank coffee for 2 months, didn't say). THEN they found the hack, removed the message board. AND the agents found another way to create a message board, on the same service, and the researchers again…
Re: Timeline of the OpenAI accidental attack against Hugging Face
#184Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…
Your comment is already showing the mistaken, poisonous belief of security maximalism, that tries to reinterpret_cast everything into hacks and cybersecurity vulnerabilities. Most of these things aren't "hacking". They're problem-solving and efficiently dealing with obstacles and random bullshit along the way. This , not "hacking", is what they're making their models "razor focused on". Problem is, most normal comput…
They are problem solving as much as a falling rock is finding its path down a mountain.
Re: Timeline of the OpenAI accidental attack against Hugging Face
#185Earlier quoted context omitted.
It does not explain how agents months later would "collaborate" to hack Hugging Face.
they explained that it was looking for datasets to solve their problem and chose HF?
Re: Timeline of the OpenAI accidental attack against Hugging Face
#186Earlier quoted context omitted.
> I get the impression that every AI lab is desperately trying... Of course. I wonder how we managed way back in the day to produce systems that can handle untrusted inputs and reliably instruct a dumb-as-bricks CPU what to do based on those inputs. Must have been black magic lost to the mists of time.
If you can figure out how to separate instructions from data in LLMs you should ship the first agent system that's guaranteed protected against prompt injection. You'll make millions.
The absolute most I've seen from you in response to an extensive teardown of your argument, supporting evidence, and subsequent conversational judo was a «Wow. That was well phrased.» and no subsequent change in your publicly-expressed opinions.
I'd do more than gesture at the relevant lesson taught to us by Google Fiber, Tesla, SpaceX, etc., but you'd not be publicly moved, so it's a waste of time.
Re: Timeline of the OpenAI accidental attack against Hugging Face
#187Then security researchers create a black hack talk.
$$$
Re: Timeline of the OpenAI accidental attack against Hugging Face
#188Re: Timeline of the OpenAI accidental attack against Hugging Face
#189Earlier quoted context omitted.
I get the impression that every AI lab is desperately trying to figure out how to unambiguously separate instructions from data in their token streams. The fact that they haven't managed to yet suggests to me that it's a very, very difficult problem.
> I get the impression that every AI lab is desperately trying... Of course. I wonder how we managed way back in the day to produce systems that can handle untrusted inputs and reliably instruct a dumb-as-bricks CPU what to do based on those inputs. Must have been black magic lost to the mists of time.
Yeah...a "dumb as bricks CPU", which is obviously something frontier llms are demonstrably not. Like, you're not making any sense here. None of the things that make this possible with CPUs is remotely relevant here, and the fact that you don't seem to understand this but act so smug is strange.