Live data from Hacker News

Path to Astra: critical capabilities and frontier safeguards

openai.com

71–80 of 110 posts

Re: Path to Astra: critical capabilities and frontier safeguards

#71
post #56

Earlier quoted context omitted.

Should everyone have access to guns too? I ask because there seems to be a disconnect where a lot of people who live in countries with gun control don’t want AI offensive capabilities to be regulated. I presume your country doesn’t have AI sovereignty so no matter what you won’t have access to models aligned with your beliefs.

> a lot of people who live in countries with gun control don’t want AI offensive capabilities to be regulated These are not really analogous: direct lethality of LLMs is quite low. Yes, they might be used to exploit a critical system, leading to loss of life, but that is a second-order effect at best. It's also not a new capability - LLMs may discover exploits faster, but cyberattacks against infrastructure targets w…

> direct lethality of LLMs is quite low.

For now, I would generally agree that it will remain low however the AI companies themselves are publically stating how capable/dangerous these models are (for whatever reason marketing/hype/upcoming IPO's, to keep the investments coming), if we take their statements at face value then prudence would suggest we stop training new ones and more fully examine the capabilities of what we've currently built.

In reality, that's not going to happen, the competition between the US and China means there will be no unilateral pause, they'll both keeping rushing/pushing as fast as they can throwing caution to the wind while doing it.

As a species though we've always done that, We just take a crowbar to Pandora's box and see what happens.

Re: Path to Astra: critical capabilities and frontier safeguards

#73
post #69

Earlier quoted context omitted.

I'm not really convinced that there's much secret sauce here, all the methods and data are public, the only real difference is how much compute it takes.

This is obviously false. There is "secret sauce" because in fact not all the methods and data are public.

Where does the secret sauce show up in the outputs, then?

Like, (apart from tone), I find it hard to distinguish between the outputs of GPT/Claude/Kimi/GLM recently (I use cursor, and have been giving them the same prompt and comparing).

If anything, I found that the non-Claude models were better in many cases, which definitely doesn't map to their pricing.

> in fact not all the methods and data are public

Probably not, but unless you work at a lab, I'm not sure that anyone can say (and if you do work at a lab, you should not be replying on this thread).

Re: Path to Astra: critical capabilities and frontier safeguards

#74
post #56

Earlier quoted context omitted.

Should everyone have access to guns too? I ask because there seems to be a disconnect where a lot of people who live in countries with gun control don’t want AI offensive capabilities to be regulated. I presume your country doesn’t have AI sovereignty so no matter what you won’t have access to models aligned with your beliefs.

> a lot of people who live in countries with gun control don’t want AI offensive capabilities to be regulated These are not really analogous: direct lethality of LLMs is quite low. Yes, they might be used to exploit a critical system, leading to loss of life, but that is a second-order effect at best. It's also not a new capability - LLMs may discover exploits faster, but cyberattacks against infrastructure targets w…

> These are not really analogous: direct lethality of LLMs is quite low.

I imagine an LLM with no safeguards and a psychopathic mind would be considerably more dangerous than a gun. I don’t mean only for hacking. Though everyone in the world suddenly having a pocket expert hacker should be taken seriously.

I find it completely insane when people bash the labs for even considering safety. The entitlement is off the charts.

Re: Path to Astra: critical capabilities and frontier safeguards

#75
post #58

Earlier quoted context omitted.

To the best of our knowledge, these Chinese companies rely on distillation of frontier models by OpenAI and Anthropic, which isn't a method available at the frontier itself.

I'm not really convinced that there's much secret sauce here, all the methods and data are public, the only real difference is how much compute it takes.

All the methods and data are not public. We don't know what unpublished methods they're using. You can get most of the pre-training data publicly but they've probably spent a ton of money curating it and are now doing things like buying rare books. The RL training data is all (/mostly) proprietary though, and that's the real secret sauce part.

Re: Path to Astra: critical capabilities and frontier safeguards

#76
post #74

Earlier quoted context omitted.

> a lot of people who live in countries with gun control don’t want AI offensive capabilities to be regulated These are not really analogous: direct lethality of LLMs is quite low. Yes, they might be used to exploit a critical system, leading to loss of life, but that is a second-order effect at best. It's also not a new capability - LLMs may discover exploits faster, but cyberattacks against infrastructure targets w…

> These are not really analogous: direct lethality of LLMs is quite low. I imagine an LLM with no safeguards and a psychopathic mind would be considerably more dangerous than a gun. I don’t mean only for hacking. Though everyone in the world suddenly having a pocket expert hacker should be taken seriously. I find it completely insane when people bash the labs for even considering safety. The entitlement is off the ch…

> I imagine an LLM with no safeguards and a psychopathic mind would be considerably more dangerous than a gun

Define the axis along which "more dangerous" is measured?

If we are scoring on lethality, guns already score 100%, instantaneous death. If we are scoring on number of casualties, guns have already been used to cause mass-fatalities.

Unless we are giving the LLM an armed drone, skynet-style, we're at best talking about indirect casualties (swatting, hacking street lights to provoke crashes, etc).

Re: Path to Astra: critical capabilities and frontier safeguards

#77
post #74

Earlier quoted context omitted.

> These are not really analogous: direct lethality of LLMs is quite low. I imagine an LLM with no safeguards and a psychopathic mind would be considerably more dangerous than a gun. I don’t mean only for hacking. Though everyone in the world suddenly having a pocket expert hacker should be taken seriously. I find it completely insane when people bash the labs for even considering safety. The entitlement is off the ch…

> I imagine an LLM with no safeguards and a psychopathic mind would be considerably more dangerous than a gun Define the axis along which "more dangerous" is measured? If we are scoring on lethality, guns already score 100%, instantaneous death. If we are scoring on number of casualties, guns have already been used to cause mass-fatalities. Unless we are giving the LLM an armed drone, skynet-style, we're at best talk…

[deleted]

Re: Path to Astra: critical capabilities and frontier safeguards

#78
post #7

> As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities. Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.

Can’t imagine the stress of the researcher who had to run exploitbench again knowing what happened last time around.

Don't worry it was subcontracted out like last time I'm sure.

Re: Path to Astra: critical capabilities and frontier safeguards

#79
I've still not seen:

* An apology for compromising a third-party's systems

* An acknowledgement of the asymmetry of defense if you're not on FrontierAI's special people list

* Anything in terms of actual safeguards that isn't "better prompt engineering"

Re: Path to Astra: critical capabilities and frontier safeguards

#80
post #75

Earlier quoted context omitted.

I'm not really convinced that there's much secret sauce here, all the methods and data are public, the only real difference is how much compute it takes.

All the methods and data are not public. We don't know what unpublished methods they're using. You can get most of the pre-training data publicly but they've probably spent a ton of money curating it and are now doing things like buying rare books. The RL training data is all (/mostly) proprietary though, and that's the real secret sauce part.

All the RL data are exactly public. There are huge amount of distilled data freely available, and that amount is more than enough to train a ~10T model.
Post reply on HN