Live data from Hacker News

Expanding Project Glasswing

anthropic.com

221–230 of 261 posts

Re: Expanding Project Glasswing

#221

Earlier quoted context omitted.

Almost all the corporate security professionals I have dealt with have been tool runners with no more than Helpdesk level skills.

As someone with over 30 years experience in computer security, both in corporate as well as boutique security and startup shops, who has been consistently fighting this trend, and recently bearing witness to and engaging in the current AI surge: I can say with absolute confidence that it is only getting and going to get even worse yet. People like me who know there is a better way are getting pushed harder to lean on…

AI is far better at security than the majority of security professionals. It is a net positive.

People constantly compare AI to this very rare expert human rather than the reality of who is already employed. Experts like you are a major culprit of this. And it puts you at odds with yourself to both admit the industry is full of subpar workers and then lament that they will be replaced with workers that are better, but still worse than you.

What is wrong with someone to make them think in this manner? Is it just a kneejerk response with little thought? Is it ego? Is it a coping mechanism? I find it very strange and interesting and annoying.

Re: Expanding Project Glasswing

#222
post #20

Is there any evidence Mythos is qualitatively better than the Opus 4.x? I'm afraid that the usual mantra that "we just need more scale" that worked well for attracting investments, is not working anymore - bigger models provide marginal improvements while naturally get much more expensive to run. Is this why both Anthropic and OpenAI are rushing for IPOs this year?

> Im afraid that the usual mantra that "we just need more scale" that worked well for attracting investments, is not working anymore - bigger models provide marginal improvements while naturally get much more expensive to run. It's super interesting to hear this refrain on HN, it is alarmingly common. Anthropic released benchmark numbers on Mythos, as they have for all of their models. Once models become public, peop…

Mythos numbers are effectively irreproducible aside from cherry-picked approvals.

Re: Expanding Project Glasswing

#225
post #106
post #40

It’s clear that Anthropic has run out of the compute capacity needed to serve Mythos publicly. They’re using security concerns to mask their inability to deliver the model at scale, while still trying to maintain their lead over OpenAI. As a result, they’ve chosen to release it privately under the banner of an “ethical” rollout.

I had to patch my Linux boxes daily at some point in the past couple months. I don’t want Mythos to be publicly released for as long as it is economically feasible for Anthropic. I hope they have a gentleman’s agreement with OpenAI and DeepMind about this, too. Chinese labs will force their hands, until then let’s hope maximum number of projects get patched at a reasonable pace.

I hope that such agreement gets broken hard and given the MSRC cold shoulder. If that means abliterated Qwen et al embarrasses Mythos to deliver a wider rollout, I’ll take that.

Trusting Anthropic to deliver is like asking Microsoft to pay out for bugs.

Re: Expanding Project Glasswing

#226
post #219

Earlier quoted context omitted.

>you can concoct plausible scenarios where it could kill every single person on the planet. Idiots can scary black box their way to that concern. Plausible? Not so much.

It is already very plausible (and has been since the 1950s) without the advent of LLMs. This is just another layer on top of the preexisting and very plausible existential threats we already face.

Detail it. Justify it.

Your comment about before LLMs is a non sequitur. Demonstrate that an LLM can kill everyone on the planet.

Re: Expanding Project Glasswing

#227

Earlier quoted context omitted.

I suppose you meant GPT-2, but for years? Did they say the same about subsequent models?

They did it for 2 and 3, however it looks like they didn't for 4 and 5. GPT-2: https://slate.com/technology/2019/02/openai-gpt2-text-genera... GPT-3: https://www.itpro.com/technology/artificial-intelligence-ai/...

For GPT-2 and GPT-3 it seems like the concern was that they hadn't yet figured out how to properly write safeguards for it yet:

> The company believes making its API generally available was made possible due to its progress with safeguards, and that opening up the API to all developers will help see applications developed faster. ...

> A large emphasis has been placed on safe use of the tool, which in the past has been criticised for a range of shortcomings, including racism and prejudices against specific genders and religions.

Re: Expanding Project Glasswing

#228

Earlier quoted context omitted.

It seems like there is a genuine communication breakdown between management and engineering. Engineers know that there are vulnerabilities all over the place and that there have been for ages and that where the rubber hits the road every vulnerability does not represent a successful exploit by some nefarious actor. Management can often treat cybersecurity like a black box that represents millions upon millions in lia…

Yeah, I’m getting the sense that Mythos is for cybersecurity what blockchain was for back-end finance. A bit useful. But mostly good for bringing attention to upgrading neglected systems.

This doesn’t make any sense either.

Many systems in relation to banking are very old and will stay that way - the economics are not favourable.

Re: Expanding Project Glasswing

#229
post #6

GPT-5.5-Cyber has already at least hit if not surpassed Mythos capability in cyber tasks. The only reason they're holding back is because once its out everyone would realize that its capabilities were a step change in March, but are not anymore, yet it costs significantly more and is much slower.

But GPT-5.5-Cyber is also not released publicly?

Re: Expanding Project Glasswing

#230

I'll share the first-hand account I recently got from someone else. > We've used it at work > it is... not as hype as everyone is concerned about > I'd argue the framework around it for security scanning is the arguably more useful side of the tool, definitely doesnt take a huge model to get all the issues it flagged on our systems > For us, it absolutely flooded us with noise > I mean hundreds if not thousands of fa…

It seems like there is a genuine communication breakdown between management and engineering. Engineers know that there are vulnerabilities all over the place and that there have been for ages and that where the rubber hits the road every vulnerability does not represent a successful exploit by some nefarious actor. Management can often treat cybersecurity like a black box that represents millions upon millions in lia…

I recommend "How to measure anything in cybersecurity risk". Really interesting read about putting actual value on security.
Post reply on HN