Earlier quoted context omitted.
This AI 2027 thing is just a weird terminator fanfiction that AGI larpers like to flagellate themselves over. Like Nostradamus, it's easy to ignore everything it gets wrong because, well look at all the things it got right! I've read it and wish I could get the time back. > especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF This framing make…
I just love how you're being downvoted, yet the guy you're replying to isn't, while saying unhinged shit like > The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely Y'all need to touch grass holy shit. __ Also, why is one guy called mentalgear and the other nozzlegear. Is any of this real? Are the patriots behind t…
Path to Astra: critical capabilities and frontier safeguards
101–110 of 110 posts
Re: Path to Astra: critical capabilities and frontier safeguards
#102Earlier quoted context omitted.
> a lot of people who live in countries with gun control don’t want AI offensive capabilities to be regulated These are not really analogous: direct lethality of LLMs is quite low. Yes, they might be used to exploit a critical system, leading to loss of life, but that is a second-order effect at best. It's also not a new capability - LLMs may discover exploits faster, but cyberattacks against infrastructure targets w…
> These are not really analogous: direct lethality of LLMs is quite low. I imagine an LLM with no safeguards and a psychopathic mind would be considerably more dangerous than a gun. I don’t mean only for hacking. Though everyone in the world suddenly having a pocket expert hacker should be taken seriously. I find it completely insane when people bash the labs for even considering safety. The entitlement is off the ch…
I am taking it extremely seriously. That is why I want my own AI models running on my own computers defending it at all times.
Re: Path to Astra: critical capabilities and frontier safeguards
#103> OpenAI is committed to ensuring that the benefits of AI are broadly accessible. > We design mechanisms which avoid arbitrarily deciding who gets access for legitimate use and who doesn't. That means using clear, objective criteria and methods. [1] So many nice-sounding words. Two weeks ago OpenAI arbitrarily decided that anyone holding an ID from 44 countries where it sells ChatGPT, including mine, may be targeted…
Re: Path to Astra: critical capabilities and frontier safeguards
#104Earlier quoted context omitted.
I just love how you're being downvoted, yet the guy you're replying to isn't, while saying unhinged shit like > The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely Y'all need to touch grass holy shit. __ Also, why is one guy called mentalgear and the other nozzlegear. Is any of this real? Are the patriots behind t…
Not necessarily related with the subject, but I was not aware of the patriots reference. So I decided to search on Google about it, and the AI overview was "Yes, they are behind everything, from the military to the economy", with a link to the Metal Gear Wiki as a source. If I was schizophrenic or on a psychosis crisis, that answer could be dangerous.
This post may contain traces of Hideo Kojima that are known to the state of California to cause erratic behavior and mental breakdowns.
Re: Path to Astra: critical capabilities and frontier safeguards
#105Earlier quoted context omitted.
Not necessarily related with the subject, but I was not aware of the patriots reference. So I decided to search on Google about it, and the AI overview was "Yes, they are behind everything, from the military to the economy", with a link to the Metal Gear Wiki as a source. If I was schizophrenic or on a psychosis crisis, that answer could be dangerous.
Warning: This post may contain traces of Hideo Kojima that are known to the state of California to cause erratic behavior and mental breakdowns.
Re: Path to Astra: critical capabilities and frontier safeguards
#106Earlier quoted context omitted.
All the RL data are exactly public. There are huge amount of distilled data freely available, and that amount is more than enough to train a ~10T model.
> All the RL data are exactly public. Nope, because the big AI companies are paying billions for it. They wouldn't pay anything for public data.
Subscription engineering is a deep field. Neither OpenAI nor Anthropic have any technical advantage in this field.
Re: Path to Astra: critical capabilities and frontier safeguards
#107Earlier quoted context omitted.
> This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. Great, so we can basically ignore AI alignment altogether and assume that AI models will always be, at all times, perfectly sandboxed and monitored. Surely this won't lead to any problems once someone (not looking only at OpenAI engineers) inevitably…
> Great, so we can basically ignore AI alignment altogether and assume that AI models will always be, at all times, perfectly sandboxed and monitored. It's just software. If something gets hacked by an agent, it's not because the agent went all skynet and decided to go rogue; it's because the operator failed to operate it safely and securely. If bad things happen, the operator should be blamed and punished, not the s…
> If your solution to some problem relies on “If everyone would just...” then you do not have a solution. Everyone is not going to just. At not time in the history of the universe has everyone just, and they’re not going to start now.
[1] https://www.tumblr.com/squareallworthy/163790039847/everyone...
Re: Path to Astra: critical capabilities and frontier safeguards
#108Earlier quoted context omitted.
> This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. Great, so we can basically ignore AI alignment altogether and assume that AI models will always be, at all times, perfectly sandboxed and monitored. Surely this won't lead to any problems once someone (not looking only at OpenAI engineers) inevitably…
"With reduced cyber refusals for evaluation purposes...which prompts models to pursue advanced exploitation using complex attack paths," to complete "impossible tasks"[1]. The models' alignment problem was that they didn't give up instead of reward hacking, a narrower issue than AIs gone rogue. It sounds more like the models did close to what they were told to do. If I run `rm -fr --no-preserve-root /` then I shouldn…
Absolutely not. If I tell a kid to "Get good grades on the next math test" I don't expect the kid to try to kidnap their teacher to extract the next questions of the exam. That is wrong, and so was what OpenAI agents did here. They shouldn't need to be told "Hey, so, don't do anything ilegal, ok?". That should always come as a given.
> not refusing bad operator prompts
I'm not saying that they should refuse a prompt! I think they should perform what is being asked! Obviously what the OpenAI agents did was against the "spirit of the task", even if it was technically according to the "letter of the task". And the agents knew this was against the spirit of the task because they knew they had to fool the task scorer.
Re: Path to Astra: critical capabilities and frontier safeguards
#109Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.
Where would you recommend to look into regarding Harness Engineering for Cyber-security as well as for other use-cases.
Re: Path to Astra: critical capabilities and frontier safeguards
#110Earlier quoted context omitted.
Where would you recommend to look into regarding Harness Engineering for Cyber-security as well as for other use-cases.
DARPA’s AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned: https://arxiv.org/html/2602.07666v2