Live data from Hacker News

Path to Astra: critical capabilities and frontier safeguards

openai.com

31–40 of 110 posts

Re: Path to Astra: critical capabilities and frontier safeguards

#32
> we believe our production safeguards at the time would have prevented the Hugging Face incident. We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions

OpenAI is fucking nuts. “Hey model you were bad last time please don’t do it again please please”.

Disconnect your training cluster from the internet for good. Physically pull the plug and only let scientists fire off experiments in the building. That’s an easy way to achieve 100% hacking protection. But I bet you that hasn’t happened and their weak sandbox will fall again…

Re: Path to Astra: critical capabilities and frontier safeguards

#33
post #27
post #22

Earlier quoted context omitted.

I've exhausted all possible avenues to get any response from OpenAI on this. Emailed them, published research that took me 2 nights to get together (saw media pick it up too), I've asked every relevant OpenAI person on X to say something, anything, saw others from Moldova also do the same. Crickets. They appear to simply not care at all.

Is there any reason why OpenAI should care? If they don't want your business then someone in your country can build a competing model.

They realistically can't. It's almost impossible to catch up to OpenAI. Only Anthropic might do it, but this is also an US American company.

Re: Path to Astra: critical capabilities and frontier safeguards

#34
post #27
post #22

Earlier quoted context omitted.

I've exhausted all possible avenues to get any response from OpenAI on this. Emailed them, published research that took me 2 nights to get together (saw media pick it up too), I've asked every relevant OpenAI person on X to say something, anything, saw others from Moldova also do the same. Crickets. They appear to simply not care at all.

Is there any reason why OpenAI should care? If they don't want your business then someone in your country can build a competing model.

They do want my business, it's an official OpenAI market. The gate isn't "we don't serve you", it's "pay for the model that may target you, but not for defensive purposes". And without a single policy document stating this, it's just a surprise gate in some random verification flow step with no explanation or appeals.

As for why they should they care, maybe they shouldn't. But then they should say which one is it, they can't care and not care at the same time.

OpenAI's own safety argument for Daybreak is that defenders need access to Critical-level models because attackers route around gates. The case for releasing Astra at all stops making sense when the gate does precisely the opposite of that.

OpenAI itself is saying plainly, that "We don’t think it’s practical or appropriate to centrally decide who gets to defend themselves. Instead, we aim to enable as many legitimate defenders as possible, with access grounded in verification, trust signals, and accountability.".

And yes, there are competing models, from China. Except OpenAI wants these models to not be accessible either.

OpenAI can pick one of two:

(a) Critical-level cyber capability is dangerous enough that access must be decided by who you are and what you do, in which case the gate has to actually look at who I am and what I do, say what the criteria are, and let me contest a wrong answer. That's their own stated policy.

(b) Access can be decided by a country code in a random dropdown, with no criteria published, no review, and no one at OpenAI able to say why - in which case drop "democratized access" and "clear, objective criteria" from the marketing, and say plainly that some passport holders don't deserve to have access to defensive capabilities.

They're now marketing a and doing b.

Re: Path to Astra: critical capabilities and frontier safeguards

#35
post #33
post #27

Earlier quoted context omitted.

Is there any reason why OpenAI should care? If they don't want your business then someone in your country can build a competing model.

They realistically can't. It's almost impossible to catch up to OpenAI. Only Anthropic might do it, but this is also an US American company.

It's not unrealistic. Several Chinese companies seem to be close behind. People thought they would never catch up to the US car industry and now look what happened.

Re: Path to Astra: critical capabilities and frontier safeguards

#36
post #27
post #22

Earlier quoted context omitted.

I've exhausted all possible avenues to get any response from OpenAI on this. Emailed them, published research that took me 2 nights to get together (saw media pick it up too), I've asked every relevant OpenAI person on X to say something, anything, saw others from Moldova also do the same. Crickets. They appear to simply not care at all.

Is there any reason why OpenAI should care? If they don't want your business then someone in your country can build a competing model.

The reason is they keep making public statements like:

> OpenAI is committed to ensuring that the benefits of AI are broadly accessible.

They can't claim that then simultaneously work to keep their cybersecurity models out of reach for non-US citizens like myself.

Re: Path to Astra: critical capabilities and frontier safeguards

#37
post #23

Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.

Where would you recommend to look into regarding Harness Engineering for Cyber-security as well as for other use-cases.

This is moreso about the (human-intended) tools, data, and environments you have available to you. Wanna do defense? Get more telemetry. Wanna do red? Get solid test-bed environments. Mature infosec programs are benefiting the most, good-guy-side wise, at the moment; because they've got these things in order already.

As far as harness engineering goes, it boils down to your ability to clearly define goals or success criteria, and safely facilitate the necessary access via the harness. There is no easy single piece of advice here, sadly. Though it would be helpful if you said what 'for Cyber-security ... other user-cases' means in your case.

Re: Path to Astra: critical capabilities and frontier safeguards

#38
post #20

They've been talking about Astra for weeks now. I wonder how much longer would they have delayed Astra, if it wasn't for Anthropic releasing Fable 5.1 today? This is why we need competition.

Meanwhile Google still hasn't released Gemini Pro 3.5

3.8 Flash tomorrow apparently, WSJ says google insiders say on-par with Opus 5. Time will tell.

Re: Path to Astra: critical capabilities and frontier safeguards

#39

I don’t see how it can be safe to release this model if it has the training history that led to the huggingface hack. You can’t just roll back that kind of reinforcement learning after the fact. Especially because these models seemed to be keenly aware that they were being evaluated by OpenAI and actively trying yo cover their tracks. How do we know that the model isn’t just pretending to be aligned?

Very simple. When the model asks to install artifactory when you give it a hard problem, you say, "no". /s

Re: Path to Astra: critical capabilities and frontier safeguards

#40
Taking as given this model meets the “Critical cybersecurity threshold” as defined by OpenAI:

Could the Federal government use the Defense Production Act or other legal tools to compel OpenAI to deliver the un-guarded model weights for national security needs?

Hard to believe any government would allow this level of capability to remain exclusively in private hands.

Interesting times.

Post reply on HN