I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra... ). The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours…
This AI 2027 thing is just a weird terminator fanfiction that AGI larpers like to flagellate themselves over. Like Nostradamus, it's easy to ignore everything it gets wrong because, well look at all the things it got right! I've read it and wish I could get the time back. > especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF This framing make…
Path to Astra: critical capabilities and frontier safeguards
41–50 of 110 posts
Re: Path to Astra: critical capabilities and frontier safeguards
#42Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.
Where would you recommend to look into regarding Harness Engineering for Cyber-security as well as for other use-cases.
Re: Path to Astra: critical capabilities and frontier safeguards
#43So the pause wasn't really a pause, got it
Re: Path to Astra: critical capabilities and frontier safeguards
#44> As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities. Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.
Can’t imagine the stress of the researcher who had to run exploitbench again knowing what happened last time around.
Re: Path to Astra: critical capabilities and frontier safeguards
#45I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra... ). The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours…
Re: Path to Astra: critical capabilities and frontier safeguards
#46Earlier quoted context omitted.
Meanwhile Google still hasn't released Gemini Pro 3.5
3.8 Flash tomorrow apparently, WSJ says google insiders say on-par with Opus 5. Time will tell.
Re: Path to Astra: critical capabilities and frontier safeguards
#47Earlier quoted context omitted.
This AI 2027 thing is just a weird terminator fanfiction that AGI larpers like to flagellate themselves over. Like Nostradamus, it's easy to ignore everything it gets wrong because, well look at all the things it got right! I've read it and wish I could get the time back. > especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF This framing make…
do you really think that a less negligent anthropic/oai/meta would really fair better against future models?
Of course, depending on which side of the terminator fanfiction you land on, you may disagree and feel that the software can rope-a-dope someone with the wherewithal to pay attention to what it's doing.
Re: Path to Astra: critical capabilities and frontier safeguards
#48> OpenAI is committed to ensuring that the benefits of AI are broadly accessible. > We design mechanisms which avoid arbitrarily deciding who gets access for legitimate use and who doesn't. That means using clear, objective criteria and methods. [1] So many nice-sounding words. Two weeks ago OpenAI arbitrarily decided that anyone holding an ID from 44 countries where it sells ChatGPT, including mine, may be targeted…
Re: Path to Astra: critical capabilities and frontier safeguards
#49They've been talking about Astra for weeks now. I wonder how much longer would they have delayed Astra, if it wasn't for Anthropic releasing Fable 5.1 today? This is why we need competition.
Re: Path to Astra: critical capabilities and frontier safeguards
#50Earlier quoted context omitted.
They realistically can't. It's almost impossible to catch up to OpenAI. Only Anthropic might do it, but this is also an US American company.
It's not unrealistic. Several Chinese companies seem to be close behind. People thought they would never catch up to the US car industry and now look what happened.