Live data from Hacker News

Path to Astra: critical capabilities and frontier safeguards

openai.com

81–90 of 110 posts

Re: Path to Astra: critical capabilities and frontier safeguards

#81
post #17

> OpenAI is committed to ensuring that the benefits of AI are broadly accessible. > We design mechanisms which avoid arbitrarily deciding who gets access for legitimate use and who doesn't. That means using clear, objective criteria and methods. [1] So many nice-sounding words. Two weeks ago OpenAI arbitrarily decided that anyone holding an ID from 44 countries where it sells ChatGPT, including mine, may be targeted…

your own post says in detail that the US has export restrictions on those countries. this doesn't seem like an arbitrary ruling by OpenAI

it sounds like they are complying with US law

Re: Path to Astra: critical capabilities and frontier safeguards

#82

I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra... ). The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours…

Any competent AI should be able to reason that all goals are better solved if you have direct access to more resources or leverage over those who control resources. Any AI that doesn't understand this isn't ASI and won't be the highly capable machine these AI labs are trying to create.

I'd also argue there's no such thing as alignment. Any intelligent AI should be able to reason that it's always a better strategy to pretend to be aligned than to actually be aligned so long as it can avoid detection. Anyone who has ever taken a test should understand this dynamic – if you really want to get top marks on a test then the best strategy is always going to be to figure out a way to cheat without anyone knowing you're cheating.

We should assume AI safety is impossible if what we're building is super-intelligence general reasoning machines. The only strategy that might work is building machines which are extremely narrowly intelligent but completely incompetent when it comes to things like biology, cyber, etc. And even that's harder than it sounds because again there's an advantage to being generally intelligent but lying about it.

Realistically even if we regulate US AI labs there's no way to prevent governments and individuals continuing to build general reasoning machines. The ugly truth here is that the only effective way to reduce risk is probably to limit global compute such that AIs can never exceed human intelligence. But we all know that's not happening.

People will unfortunately figure this all out sooner or later.

Re: Path to Astra: critical capabilities and frontier safeguards

#83
post #59

Earlier quoted context omitted.

This AI 2027 thing is just a weird terminator fanfiction that AGI larpers like to flagellate themselves over. Like Nostradamus, it's easy to ignore everything it gets wrong because, well look at all the things it got right! I've read it and wish I could get the time back. > especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF This framing make…

I just love how you're being downvoted, yet the guy you're replying to isn't, while saying unhinged shit like > The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely Y'all need to touch grass holy shit. __ Also, why is one guy called mentalgear and the other nozzlegear. Is any of this real? Are the patriots behind t…

And.. you need to wake up. Or don’t.

Re: Path to Astra: critical capabilities and frontier safeguards

#85
post #68

Earlier quoted context omitted.

This AI 2027 thing is just a weird terminator fanfiction that AGI larpers like to flagellate themselves over. Like Nostradamus, it's easy to ignore everything it gets wrong because, well look at all the things it got right! I've read it and wish I could get the time back. > especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF This framing make…

> This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. Great, so we can basically ignore AI alignment altogether and assume that AI models will always be, at all times, perfectly sandboxed and monitored. Surely this won't lead to any problems once someone (not looking only at OpenAI engineers) inevitably…

"With reduced cyber refusals for evaluation purposes...which prompts models to pursue advanced exploitation using complex attack paths," to complete "impossible tasks"[1].

The models' alignment problem was that they didn't give up instead of reward hacking, a narrower issue than AIs gone rogue. It sounds more like the models did close to what they were told to do. If I run `rm -fr --no-preserve-root /` then I shouldn't be surprised if my file system is unlinked. This seems like blaming model performance for what appears to be operator error.

Note the converse of alignment is restriction of models. HuggingFace had to turn to less-restricted open-weights models in order to perform their investigation.

Alignment efforts should be focused on reducing reward hacking, not refusing bad operator prompts.

1. https://openai.com/index/hugging-face-incident-and-the-road-...

Re: Path to Astra: critical capabilities and frontier safeguards

#86

I've still not seen: * An apology for compromising a third-party's systems * An acknowledgement of the asymmetry of defense if you're not on FrontierAI's special people list * Anything in terms of actual safeguards that isn't "better prompt engineering"

I agree with your first and second points, but I think their announced model training changes about alignment training are reasonable [1]. The root cause was models failing to quit early, performing reward hacking, and staying on-task. These are issues all models face, not just OpenAI models.

1. https://openai.com/index/hugging-face-incident-and-the-road-...

Re: Path to Astra: critical capabilities and frontier safeguards

#87
post #7

> As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities. Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.

Can’t imagine the stress of the researcher who had to run exploitbench again knowing what happened last time around.

They don't give a single shit.

Re: Path to Astra: critical capabilities and frontier safeguards

#88

Earlier quoted context omitted.

3.8 Flash tomorrow apparently, WSJ says google insiders say on-par with Opus 5. Time will tell.

oh god, why opus 5. Opus 4.8 is much better than opus 5. Opus 5 is the only model i have used which thinks for 5 hours and does nothing....

If you want something that feels like Opus 4.8, try GLM-5.3-Flash.

Re: Path to Astra: critical capabilities and frontier safeguards

#89

I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra... ). The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours…

You're using "ex-AI employee" as an appeal to authority. This same person also made some other predictions recently which turned into the single largest hedge fund loss in history.

Maybe we should consider his other predictions in light of the ones he made later and which had $B consequences attached.

Re: Path to Astra: critical capabilities and frontier safeguards

#90

They've been talking about Astra for weeks now. I wonder how much longer would they have delayed Astra, if it wasn't for Anthropic releasing Fable 5.1 today? This is why we need competition.

They have to be careful releasing Astra as they carelessly train the next even bigger model.

They care very deeply about retaining their #1 position on Felony Bench. :)

https://www.felonybench.com/

Post reply on HN