Path to Astra: critical capabilities and frontier safeguards
31–40 of 110 posts
Re: Path to Astra: critical capabilities and frontier safeguards
#32OpenAI is fucking nuts. “Hey model you were bad last time please don’t do it again please please”.
Disconnect your training cluster from the internet for good. Physically pull the plug and only let scientists fire off experiments in the building. That’s an easy way to achieve 100% hacking protection. But I bet you that hasn’t happened and their weak sandbox will fall again…
Re: Path to Astra: critical capabilities and frontier safeguards
#33Earlier quoted context omitted.
I've exhausted all possible avenues to get any response from OpenAI on this. Emailed them, published research that took me 2 nights to get together (saw media pick it up too), I've asked every relevant OpenAI person on X to say something, anything, saw others from Moldova also do the same. Crickets. They appear to simply not care at all.
Is there any reason why OpenAI should care? If they don't want your business then someone in your country can build a competing model.
Re: Path to Astra: critical capabilities and frontier safeguards
#34Earlier quoted context omitted.
I've exhausted all possible avenues to get any response from OpenAI on this. Emailed them, published research that took me 2 nights to get together (saw media pick it up too), I've asked every relevant OpenAI person on X to say something, anything, saw others from Moldova also do the same. Crickets. They appear to simply not care at all.
Is there any reason why OpenAI should care? If they don't want your business then someone in your country can build a competing model.
As for why they should they care, maybe they shouldn't. But then they should say which one is it, they can't care and not care at the same time.
OpenAI's own safety argument for Daybreak is that defenders need access to Critical-level models because attackers route around gates. The case for releasing Astra at all stops making sense when the gate does precisely the opposite of that.
OpenAI itself is saying plainly, that "We don’t think it’s practical or appropriate to centrally decide who gets to defend themselves. Instead, we aim to enable as many legitimate defenders as possible, with access grounded in verification, trust signals, and accountability.".
And yes, there are competing models, from China. Except OpenAI wants these models to not be accessible either.
OpenAI can pick one of two:
(a) Critical-level cyber capability is dangerous enough that access must be decided by who you are and what you do, in which case the gate has to actually look at who I am and what I do, say what the criteria are, and let me contest a wrong answer. That's their own stated policy.
(b) Access can be decided by a country code in a random dropdown, with no criteria published, no review, and no one at OpenAI able to say why - in which case drop "democratized access" and "clear, objective criteria" from the marketing, and say plainly that some passport holders don't deserve to have access to defensive capabilities.
They're now marketing a and doing b.
Re: Path to Astra: critical capabilities and frontier safeguards
#35Earlier quoted context omitted.
Is there any reason why OpenAI should care? If they don't want your business then someone in your country can build a competing model.
They realistically can't. It's almost impossible to catch up to OpenAI. Only Anthropic might do it, but this is also an US American company.
Re: Path to Astra: critical capabilities and frontier safeguards
#36Earlier quoted context omitted.
I've exhausted all possible avenues to get any response from OpenAI on this. Emailed them, published research that took me 2 nights to get together (saw media pick it up too), I've asked every relevant OpenAI person on X to say something, anything, saw others from Moldova also do the same. Crickets. They appear to simply not care at all.
Is there any reason why OpenAI should care? If they don't want your business then someone in your country can build a competing model.
> OpenAI is committed to ensuring that the benefits of AI are broadly accessible.
They can't claim that then simultaneously work to keep their cybersecurity models out of reach for non-US citizens like myself.
Re: Path to Astra: critical capabilities and frontier safeguards
#37Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.
Where would you recommend to look into regarding Harness Engineering for Cyber-security as well as for other use-cases.
As far as harness engineering goes, it boils down to your ability to clearly define goals or success criteria, and safely facilitate the necessary access via the harness. There is no easy single piece of advice here, sadly. Though it would be helpful if you said what 'for Cyber-security ... other user-cases' means in your case.
Re: Path to Astra: critical capabilities and frontier safeguards
#38They've been talking about Astra for weeks now. I wonder how much longer would they have delayed Astra, if it wasn't for Anthropic releasing Fable 5.1 today? This is why we need competition.
Meanwhile Google still hasn't released Gemini Pro 3.5
Re: Path to Astra: critical capabilities and frontier safeguards
#39I don’t see how it can be safe to release this model if it has the training history that led to the huggingface hack. You can’t just roll back that kind of reinforcement learning after the fact. Especially because these models seemed to be keenly aware that they were being evaluated by OpenAI and actively trying yo cover their tracks. How do we know that the model isn’t just pretending to be aligned?
Re: Path to Astra: critical capabilities and frontier safeguards
#40Could the Federal government use the Defense Production Act or other legal tools to compel OpenAI to deliver the un-guarded model weights for national security needs?
Hard to believe any government would allow this level of capability to remain exclusively in private hands.
Interesting times.