Live data from Hacker News

Path to Astra: critical capabilities and frontier safeguards

openai.com

101–110 of 110 posts

Re: Path to Astra: critical capabilities and frontier safeguards

#101
post #59

Earlier quoted context omitted.

This AI 2027 thing is just a weird terminator fanfiction that AGI larpers like to flagellate themselves over. Like Nostradamus, it's easy to ignore everything it gets wrong because, well look at all the things it got right! I've read it and wish I could get the time back. > especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF This framing make…

I just love how you're being downvoted, yet the guy you're replying to isn't, while saying unhinged shit like > The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely Y'all need to touch grass holy shit. __ Also, why is one guy called mentalgear and the other nozzlegear. Is any of this real? Are the patriots behind t…

Not necessarily related with the subject, but I was not aware of the patriots reference. So I decided to search on Google about it, and the AI overview was "Yes, they are behind everything, from the military to the economy", with a link to the Metal Gear Wiki as a source. If I was schizophrenic or on a psychosis crisis, that answer could be dangerous.

Re: Path to Astra: critical capabilities and frontier safeguards

#102
post #74

Earlier quoted context omitted.

> a lot of people who live in countries with gun control don’t want AI offensive capabilities to be regulated These are not really analogous: direct lethality of LLMs is quite low. Yes, they might be used to exploit a critical system, leading to loss of life, but that is a second-order effect at best. It's also not a new capability - LLMs may discover exploits faster, but cyberattacks against infrastructure targets w…

> These are not really analogous: direct lethality of LLMs is quite low. I imagine an LLM with no safeguards and a psychopathic mind would be considerably more dangerous than a gun. I don’t mean only for hacking. Though everyone in the world suddenly having a pocket expert hacker should be taken seriously. I find it completely insane when people bash the labs for even considering safety. The entitlement is off the ch…

> Though everyone in the world suddenly having a pocket expert hacker should be taken seriously.

I am taking it extremely seriously. That is why I want my own AI models running on my own computers defending it at all times.

Re: Path to Astra: critical capabilities and frontier safeguards

#103
post #17

> OpenAI is committed to ensuring that the benefits of AI are broadly accessible. > We design mechanisms which avoid arbitrarily deciding who gets access for legitimate use and who doesn't. That means using clear, objective criteria and methods. [1] So many nice-sounding words. Two weeks ago OpenAI arbitrarily decided that anyone holding an ID from 44 countries where it sells ChatGPT, including mine, may be targeted…

I'm so confused by the point you're trying to make. There's a lot of rhetoric about Georgia and an investigation about how the country list is 1:1 with some mysterious list from 1996 - but like - it's an export control list? Yeah, Washington approved the sale of missiles to Georgia. That's how that list works. Washington has to approve it. Openai is not Washington. They're perfectly reasonably erring on the side of caution and potentially over-complying with export controls. And if they then have to get approval from the feds to export to Georgia - well - our current administration has provided many reasons to over comply with trade related controls and not exactly been a champion of encouraging cross country trade right now. Idk what you expect from OpenAI or why you think it would matter at all that Georgia is a democracy or an ally of the US or is closely aligned with us. Ask Canada and NATO how much that's counted for with this administration.

Re: Path to Astra: critical capabilities and frontier safeguards

#104
post #59

Earlier quoted context omitted.

I just love how you're being downvoted, yet the guy you're replying to isn't, while saying unhinged shit like > The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely Y'all need to touch grass holy shit. __ Also, why is one guy called mentalgear and the other nozzlegear. Is any of this real? Are the patriots behind t…

Not necessarily related with the subject, but I was not aware of the patriots reference. So I decided to search on Google about it, and the AI overview was "Yes, they are behind everything, from the military to the economy", with a link to the Metal Gear Wiki as a source. If I was schizophrenic or on a psychosis crisis, that answer could be dangerous.

Warning:

This post may contain traces of Hideo Kojima that are known to the state of California to cause erratic behavior and mental breakdowns.

Re: Path to Astra: critical capabilities and frontier safeguards

#105
post #104

Earlier quoted context omitted.

Not necessarily related with the subject, but I was not aware of the patriots reference. So I decided to search on Google about it, and the AI overview was "Yes, they are behind everything, from the military to the economy", with a link to the Metal Gear Wiki as a source. If I was schizophrenic or on a psychosis crisis, that answer could be dangerous.

Warning: This post may contain traces of Hideo Kojima that are known to the state of California to cause erratic behavior and mental breakdowns.

I'm sorry if I made it sound like I was accusing you of something. My comment was about the AI overview result. It didn't explain what is the reference, just giving what I said in the previous comment. I was confused until I hovered in the link in the end of the overview card

Re: Path to Astra: critical capabilities and frontier safeguards

#106
post #96

Earlier quoted context omitted.

All the RL data are exactly public. There are huge amount of distilled data freely available, and that amount is more than enough to train a ~10T model.

> All the RL data are exactly public. Nope, because the big AI companies are paying billions for it. They wouldn't pay anything for public data.

There are 'transfer stations' and that's how exactly I use GPT and Claude in China. OpenAI and Anthropic do not sell in China, so we use their AI with a much lower price like 1% of the official API price. The largest transfer stations have TBs of traffic every day, and the traffic is eventually possessed by the open source community.

Subscription engineering is a deep field. Neither OpenAI nor Anthropic have any technical advantage in this field.

Re: Path to Astra: critical capabilities and frontier safeguards

#107
post #68

Earlier quoted context omitted.

> This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. Great, so we can basically ignore AI alignment altogether and assume that AI models will always be, at all times, perfectly sandboxed and monitored. Surely this won't lead to any problems once someone (not looking only at OpenAI engineers) inevitably…

> Great, so we can basically ignore AI alignment altogether and assume that AI models will always be, at all times, perfectly sandboxed and monitored. It's just software. If something gets hacked by an agent, it's not because the agent went all skynet and decided to go rogue; it's because the operator failed to operate it safely and securely. If bad things happen, the operator should be blamed and punished, not the s…

I don't want someone to blame. I want agents to be aligned by default. Their good behavior shouldn't depend on all users at all times using them correctly, because everyone will not just[1] use them correctly at all times.

> If your solution to some problem relies on “If everyone would just...” then you do not have a solution. Everyone is not going to just. At not time in the history of the universe has everyone just, and they’re not going to start now.

[1] https://www.tumblr.com/squareallworthy/163790039847/everyone...

Re: Path to Astra: critical capabilities and frontier safeguards

#108
post #68

Earlier quoted context omitted.

> This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. Great, so we can basically ignore AI alignment altogether and assume that AI models will always be, at all times, perfectly sandboxed and monitored. Surely this won't lead to any problems once someone (not looking only at OpenAI engineers) inevitably…

"With reduced cyber refusals for evaluation purposes...which prompts models to pursue advanced exploitation using complex attack paths," to complete "impossible tasks"[1]. The models' alignment problem was that they didn't give up instead of reward hacking, a narrower issue than AIs gone rogue. It sounds more like the models did close to what they were told to do. If I run `rm -fr --no-preserve-root /` then I shouldn…

> It sounds more like the models did close to what they were told to do

Absolutely not. If I tell a kid to "Get good grades on the next math test" I don't expect the kid to try to kidnap their teacher to extract the next questions of the exam. That is wrong, and so was what OpenAI agents did here. They shouldn't need to be told "Hey, so, don't do anything ilegal, ok?". That should always come as a given.

> not refusing bad operator prompts

I'm not saying that they should refuse a prompt! I think they should perform what is being asked! Obviously what the OpenAI agents did was against the "spirit of the task", even if it was technically according to the "letter of the task". And the agents knew this was against the spirit of the task because they knew they had to fool the task scorer.

Re: Path to Astra: critical capabilities and frontier safeguards

#109
post #23

Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.

Where would you recommend to look into regarding Harness Engineering for Cyber-security as well as for other use-cases.

We have posted some articles on it (blog.vulnetic.ai), Im also happy to chat!

Re: Path to Astra: critical capabilities and frontier safeguards

#110
post #42
post #23

Earlier quoted context omitted.

Where would you recommend to look into regarding Harness Engineering for Cyber-security as well as for other use-cases.

DARPA’s AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned: https://arxiv.org/html/2602.07666v2

this might be dated. the tech moves very fast in this space and architecture from 6 months ago is dated.
Post reply on HN