Live data from Hacker News

Expanding Project Glasswing

anthropic.com

241–250 of 261 posts

Re: Expanding Project Glasswing

#241

Earlier quoted context omitted.

They did it for 2 and 3, however it looks like they didn't for 4 and 5. GPT-2: https://slate.com/technology/2019/02/openai-gpt2-text-genera... GPT-3: https://www.itpro.com/technology/artificial-intelligence-ai/...

For GPT-2 and GPT-3 it seems like the concern was that they hadn't yet figured out how to properly write safeguards for it yet: > The company believes making its API generally available was made possible due to its progress with safeguards, and that opening up the API to all developers will help see applications developed faster. ... > A large emphasis has been placed on safe use of the tool, which in the past has be…

Maybe, but they certainly used it for marketing too. At the time they contacted a bunch of publications and gave them access but told them they could only share snippets of the output [1]. The only reason to set restrictions like that is marketing.

[1]: https://youtu.be/TfVYxnhuEdU?t=102

Transcript of the timestamped part:

> Now, OpenAI's terms of service don't let me give you the full list. I have to curate them, and show you a sample. Those are the terms and conditions I agreed to.

Re: Expanding Project Glasswing

#242
post #236

Earlier quoted context omitted.

You're kind of an asshole. No thanks

Its not like you handed me anything but woo to work with. There's really nothing less respectful than making up absolute nonsense and expecting a kind and thoughtful reply.

No they're right. Regardless if one agrees with you or not, doesn't change the fact that your behavior was that of an asshole. I would know since I'm one too.

Re: Expanding Project Glasswing

#243
post #221

Earlier quoted context omitted.

AI is far better at security than the majority of security professionals. It is a net positive. People constantly compare AI to this very rare expert human rather than the reality of who is already employed. Experts like you are a major culprit of this. And it puts you at odds with yourself to both admit the industry is full of subpar workers and then lament that they will be replaced with workers that are better, bu…

I also don’t like your framing, here. We need experts to know when AI is wrong, which it is all the time. Earlier this week someone commented here that we shouldn’t expect a language model to know that you need to drive a car to a car wash, to wash a car. So then, what do we expect it to know? Who’s responsible for when it’s wrong? Also, why can’t Mythos just fix all these issues itself if it’s so smart. And test the…

> why can’t Mythos just fix all these issues itself if it’s so smart. And test them to make sure they work?

“Why”: because you didn’t ask it. It’s not its job in this case.

You don’t hire an accountant and tell them “why can’t you fix my cash-flow problems and make me money if you’re so smart”

Re: Expanding Project Glasswing

#244

Earlier quoted context omitted.

why are folks looking at the output of the first pass? my understanding, and experience, is that you 1. run a bunch of sessions with small permutations to create variety , 2. run more sessions dedupe reports into a smaller collections of potential vulns, 3. run a handful of agents at max effort to write PoCs + write-ups, 4. rank findings, 5. finally look at what, if anything that, was found. maybe ask questions, try…

What Project Glasswing claimed at launch is that Mythos can "surpass all but the most skilled humans at finding and exploiting software vulnerabilities". What you're describing sounds more like making skilled humans more effective at penetration testing. That's cool, but it's not clear how much it matters, because most security teams were not previously bottlenecked on penetration testing capacity.

i wasn't thinking about pen-testing, but vulnerability-research, which seems to match that quote. but, you're right, gp is referring to "security scanning". i just feel like, even then whoever's running the research, should triage and validate results, before passing on to mgmt.

Re: Expanding Project Glasswing

#245
post #224
post #21

Earlier quoted context omitted.

Why would it not be reproducible?

Its analysis from prompt/harness to end products.

You can't do that for me either.

Why would you refuse to use a patch that deals with a valid PoC exploit?

If a random contributor posted an explanation of an exploit, showed it worked in an executable way, presented a patch and you could see that the exploit no longer worked - would you refuse to use the fix until the contributor showed how they figured it out?

Re: Expanding Project Glasswing

#246

Earlier quoted context omitted.

Jack Clark, co-founder of Anthropic said the following at an Oxford lecture last week ([0], at around 10 and 12 mins): "It's a technology that we do not fully understand because it's more grown than made. And it is a technology that you can concoct plausible scenarios where it could kill every single person on the planet. So to think building this technology is without risk would be an act of hubris or insanity. [...…

Thank you Mr. Altman for firing the starting gun when no one else wanted to race. (The ambiguity of sarcasm is intentional here.)

Didnt Google start the race with their paper?

Re: Expanding Project Glasswing

#247
post #245
post #224

Earlier quoted context omitted.

Its analysis from prompt/harness to end products.

You can't do that for me either. Why would you refuse to use a patch that deals with a valid PoC exploit? If a random contributor posted an explanation of an exploit, showed it worked in an executable way, presented a patch and you could see that the exploit no longer worked - would you refuse to use the fix until the contributor showed how they figured it out?

Given where Mythos alleges to go, reproducibility far beyond a hash promise, an alleged (but not really proven) existence of an PoC, and “Trust me bro” is necessary.

When an ungated (or even abliterated) public model can repeatedly, easily, and accurately embarrass Anthropic’s models, that might change.

Re: Expanding Project Glasswing

#248
post #246

Earlier quoted context omitted.

Thank you Mr. Altman for firing the starting gun when no one else wanted to race. (The ambiguity of sarcasm is intentional here.)

Didnt Google start the race with their paper?

Google (and other labs) wanted to keep the tech internal because of the obvious safety concerns. Once they were confident that the tech was understood and under control, the public could start being drip fed. Everyone on the ground back then was hyper cautious.

Then Altman made ChatGPT public, and the race began.

Re: Expanding Project Glasswing

#249

Earlier quoted context omitted.

I also don’t like your framing, here. We need experts to know when AI is wrong, which it is all the time. Earlier this week someone commented here that we shouldn’t expect a language model to know that you need to drive a car to a car wash, to wash a car. So then, what do we expect it to know? Who’s responsible for when it’s wrong? Also, why can’t Mythos just fix all these issues itself if it’s so smart. And test the…

> why can’t Mythos just fix all these issues itself if it’s so smart. And test them to make sure they work? “Why”: because you didn’t ask it. It’s not its job in this case. You don’t hire an accountant and tell them “why can’t you fix my cash-flow problems and make me money if you’re so smart”

Ah ok, sure. The difference being the model should know how to do both based on what I’ve been told.

So why didn’t Anthropic ask it for me?

Re: Expanding Project Glasswing

#250

Earlier quoted context omitted.

Have you read firefoxes findings? They found it to be qualitatively improved over Opus, and have published several of the resulting CVEs as well as more detailed numbers.

They also seem to point to it being more the harness than the model itself.

Really? They mention that Opus 4.7 in the same harness found like 1% of the bugs that Mythos found.
Post reply on HN