Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…
The Guardian's article and your reply here are so foolish and absurd that I can only imagine OpenAI employees are cringing but know they can't/shouldn't really say much.
Be skeptical of OpenAI's rogue hacker agent story
91–100 of 321 posts
Re: Be skeptical of OpenAI's rogue hacker agent story
#92Earlier quoted context omitted.
The Guardian's article and your reply here are so foolish and absurd that I can only imagine OpenAI employees are cringing but know they can't/shouldn't really say much.
touched a nerve huh?
Re: Be skeptical of OpenAI's rogue hacker agent story
#93Crazy that we need reminders not to take everything we read in corporate press releases and marketing material at face value
Re: Be skeptical of OpenAI's rogue hacker agent story
#94"It's a marketing stunt" is just denial trying to look like it's being clever.
Re: Be skeptical of OpenAI's rogue hacker agent story
#95> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit. It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that…
I agree with you that the article was disappointingly light, but there is not in general a singular correct answer when it comes to interpreting narrative and I don't think that is a useful way to evaluate what the author has written, nor real-world writing in general. The author does not implore you to have a singular 'correct' point of view on the matter but rather to hesitate before repeating the claim that "the A…
Re: Be skeptical of OpenAI's rogue hacker agent story
#96Re: Be skeptical of OpenAI's rogue hacker agent story
#97Earlier quoted context omitted.
Then train your agent on the Bible. Honestly, all the agent did was duplicate human behaviour and that better than the human. The agent was trained on human data and did what any other human would have done. To believe that agents will inherently be morally better than us is an illusion - sorry to say but that's the case. The alternative would be that the AI is truly conscious and can reason that it won't behave as i…
Dude.. what are you smoking..? It’s not about morality. It’s about asking it to do task A and doing task B with the hope of getting the result of task A as a byproduct. Meaning you will have to spend more time and tokens to actually get it to do what you want it to do. What do the bible and morals have anything to do with it? I’m criticizing the behavior I see even in the current models. You ask it to do A and instea…
The only way it would decide to do this is prompting with a deliberate combination of omissions and reiterating that the only thing that matters is the end score regardless of method.
Re: Be skeptical of OpenAI's rogue hacker agent story
#98Earlier quoted context omitted.
Despite the common misconceptions from TV, the victim "pressing charges" isn't actually a thing in criminal cases: prosecutors can choose to put someone on trial even if the victim doesn't want that. In practice this is somewhat rare, but it certainly can happen. In my reply to tokioyoyo below I laid out why this is one instance where the government should prosecute even if HuggingFace doesn't want it to.
Criminal Cases of 'hacking' require specific intent. What you're asking is that the prosecution attempt to prove Open AI intended to infiltrate Huggingface maliciously, all while the victim is saying 'no harm no foul'. No offense but prosecutors have better things to do with their time.
Re: Be skeptical of OpenAI's rogue hacker agent story
#99> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit. It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that…
Re: Be skeptical of OpenAI's rogue hacker agent story
#100Earlier quoted context omitted.
> Agents don't work on their own This is factually false: they both can and clearly did operate in an autonomous and unsupervised manner: https://openai.com/index/hugging-face-model-evaluation-secur... This does not require sentience, personhood, a soul, or anything of the sort. It further doesn't mean an erasure of legal responsibility, not in principle, and not in historical practice. I wish people would finally st…
>> Agents don't work on their own > This is factually false From your link: > After investigating, we now know that this particular incident was driven by a combination of OpenAI models...while being internally tested on a benchmark of cyber capabilities. Someone set up that test and started it. Whether they outsourced the majority of the work in "setting up" and "starting it" to an LLM or not, they still set it in m…
There's no indication of there having been a human in the loop during its operation: nobody was approving its tool calls, and nobody instructed it to commit these specific actions during its run (via prompting or steering).
There's no indication of any supervision of its operation either: OpenAI's engineers acted with significant delay, long after the agent has already meandered its way through their own infrastructure first.
Given that setting up this contraption in an insufficiently secure manner is almost certainly already a legal liability of equal significance, rejecting this very clear structural distinction is not necessary. That is unless someone is biased towards not wanting to grant the label of autonomy to it, in which case yes, this is absolutely spiritualistic reasoning, hence my point.
I do not want regulation to ride on people's nebulous identification on what specific traits and labels count as human-exclusive. Not just because I deeply disagree that e.g. autonomy would [0], but also because it is entirely unnecessary, for the reasons you also lay out. The agent having operated autonomously doesn't wash OpenAI of responsibility - so why reject the label, if not on a spiritualistic basis?
[0] thousands of years old idea that it is not, by the way: https://en.wikipedia.org/wiki/Automaton -- see also existing regulation recognizing this idea and working with it fine
Edit: one might also want to consider if the law should bite different if there was a human in the loop, or if there were explicit instructions for the agent to take unlawful actions. I'd say yes, and then that also requires this distinction to exist.