Live data from Hacker News

Path to Astra: critical capabilities and frontier safeguards

openai.com

1–10 of 110 posts

Re: Path to Astra: critical capabilities and frontier safeguards

#4
It's been a busy month at OpenAI.

I'm looking forward to seeing the increased coordination and engineering skills from Astra - one of the charts shows it roughly 2-3x better in 50% of the tokens from 5.6 sol, which I find to be very capable, if still a bit 'linearly minded' when given instructions. Even in fast mode, I wish sol were quicker, so token efficiency is greatly appreciated.

Adding these cyber capabilities has let me do a bunch of low grade IT tasks around my house I've been putting off, like updating an old home assistant raspberry pi, and one way to use the cyber capacity for good is liberating (and keeping free) weird cloud hardware we have floating around the house, so I'm hoping for some nice dividends in terms of true ownership of hardware we've got.

Re: Path to Astra: critical capabilities and frontier safeguards

#5
> As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities.

Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.

Re: Path to Astra: critical capabilities and frontier safeguards

#6
post #4

It's been a busy month at OpenAI. I'm looking forward to seeing the increased coordination and engineering skills from Astra - one of the charts shows it roughly 2-3x better in 50% of the tokens from 5.6 sol, which I find to be very capable, if still a bit 'linearly minded' when given instructions. Even in fast mode, I wish sol were quicker, so token efficiency is greatly appreciated. Adding these cyber capabilities…

I am interested in seeing how much these cybersecurity capabilities correlate to general programming. Cybersecurity definitely feels like it would be easier for an agent due to the natural explicit feedback "did I get access or not". While general programming has many less-explicit concerns (is the code readable/maintainable, robust, bug-free, performant, scalable etc).

Re: Path to Astra: critical capabilities and frontier safeguards

#7

> As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities. Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.

Can’t imagine the stress of the researcher who had to run exploitbench again knowing what happened last time around.

Re: Path to Astra: critical capabilities and frontier safeguards

#8
I don’t see how it can be safe to release this model if it has the training history that led to the huggingface hack. You can’t just roll back that kind of reinforcement learning after the fact.

Especially because these models seemed to be keenly aware that they were being evaluated by OpenAI and actively trying yo cover their tracks. How do we know that the model isn’t just pretending to be aligned?

Re: Path to Astra: critical capabilities and frontier safeguards

#9
I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra...).

The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely with Altman's golden marketing-hype boy leadership pushing the for-profit gas pedal like this.

Honestly, this is just pure irresponsible insanity to play with the fate of the world - basically a death race of the biggest few tech companies on the planet. And if you think I'm being dramatic, listen in again to ex oAI employee[0] and check for yourself how chillingly on trajectory we already are.

[0] https://ai-2027.com/

Re: Path to Astra: critical capabilities and frontier safeguards

#10
post #7

> As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities. Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.

Can’t imagine the stress of the researcher who had to run exploitbench again knowing what happened last time around.

Could be risky. Yet goal solution.
Post reply on HN