Live data from Hacker News

Expanding Project Glasswing

anthropic.com

71–80 of 261 posts

Re: Expanding Project Glasswing

#71

Maybe it is just me: I feel Anthropic most recent product announcements resemble more and more like what IBM tactic was at its high. For instance, the Watson AI hype after it defeated Kasparov. The difference is IBM actually wanted and let businesses buy and use Watson as opposed to time released like what Anthropic does to even boost the hype higher.

Big Blue defeated Kasparov. The Watson hype was about winning Jeopardy, which is still kind of the only use case for current AI.

Re: Expanding Project Glasswing

#72

Step 1: claim you created a tool so dangerous you can't release it Step2: offer to test it, but only for the biggest companies in the world Step 3: onboard those big players on your tooling and product Step 4: profit This is genius.

It’s true that providing security services to so many organizations will likely put them in a position to earn lots of money. It makes them an essential service, sort of like what happened with Cloudflare and denial-of-service attacks. (There are competitors, but they’re the first company people think of.) But I think that downplays the importance of having a good product. If the product didn’t work, this would be a…

This is a circular economy that makes everyone look good. Almost all of these enterprise companies are sitting on top of so much of tech debt that in any realistic scenario they cant really patch vulnerabilities if they are even in double digits. A lot of these companies would not even let their valuable enough code to be ingested by LLM's.

At this phase no company would risk their brand by calling the product as ineffective. The big players are in it together and small ones have no option but to play along.

Nevertheless collecting the historical wisdom and running it at machine scale does have a lot of benefits for sure. The only question is the signal to noise ratio, machine is doing what humans did, just at a multiplier speed and with a lot more context than what a normal human can hold.

Re: Expanding Project Glasswing

#73

Step 1: claim you created a tool so dangerous you can't release it Step2: offer to test it, but only for the biggest companies in the world Step 3: onboard those big players on your tooling and product Step 4: profit This is genius.

And put Chris Olah, Anthropic co-founder, sitting next to Pope Leo XIV presenting his first encyclical, Magnifica Humanitas, at the Vatican.

Re: Expanding Project Glasswing

#74

Earlier quoted context omitted.

The brits have a step-based benchmark that they use for this - https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5... They seem pretty close, in both average and "best run" scores. And, in a highly verifiable domain, "best run" or pass@n is what you're looking for.

Worth looking at the followup post that evaluates the current version of Mythos, which solves one of the main tasks that GPT-5.5-Cyber does not. https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber...

I believe the correct way to interpret AISI’s findings is that both Mythos and 5.5-Cyber are capable of solving their full benchmark (the only two models that can); Mythos does it with fewer tokens and more consistently.

Two things of note: 5.5-Cyber is likely to be substantially cheaper than Mythos, given it is priced around Opus. Additionally: AISI has never tested OpenAI’s best public model and actual Mythos competitor: 5.5-Pro.

Re: Expanding Project Glasswing

#75
post #40

It’s clear that Anthropic has run out of the compute capacity needed to serve Mythos publicly. They’re using security concerns to mask their inability to deliver the model at scale, while still trying to maintain their lead over OpenAI. As a result, they’ve chosen to release it privately under the banner of an “ethical” rollout.

It is not "clear", as your comment suggests, it's hidden. Which is semantically the opposite of clear. Regarding your theory, might be true, might be false. But it's highly speculative.

All of us, including you, know that he is not saying "they are being transparent." When someone says "it's clear that..." in this way they're saying "It's clear to us what is really happening here.

Re: Expanding Project Glasswing

#76
post #33

In case the topic of memory safety is interesting to anyone I've been experimenting with using AI agents to port common web infra projects to safe/ performant Rust. Somewhat inspired by the Bun port - was thinking that at some point memory safety might be such a big deal that people just need drop in replacements. - Valkey/ Redis port here https://github.com/ianm199/valdr (passes ~99% of single node test suite, real…

[dead]

Re: Expanding Project Glasswing

#77

Step 1: claim you created a tool so dangerous you can't release it Step2: offer to test it, but only for the biggest companies in the world Step 3: onboard those big players on your tooling and product Step 4: profit This is genius.

Seems like they're not even close to step 4.

Re: Expanding Project Glasswing

#78
post #40

It’s clear that Anthropic has run out of the compute capacity needed to serve Mythos publicly. They’re using security concerns to mask their inability to deliver the model at scale, while still trying to maintain their lead over OpenAI. As a result, they’ve chosen to release it privately under the banner of an “ethical” rollout.

I find this line of reasoning highly dubious.

Yes, Anthropic is compute constrained, even after the SpaceX Colossus deal.

But supply constraints are the normal operating mode of any market. Anthropic could choose to serve whatever models it pleases at whatever price points it chooses and let the market decide where the value is.

If Mythos at $X overwhelms their capacity, they could just charge $X+1. If still overwhelmed, there are larger prices as well.

Re: Expanding Project Glasswing

#79

Earlier quoted context omitted.

Where did I use the word scam? Marketing move doesn't mean scam. It describe the ability to sell people over a narrative and surpassing your competitor in market share. And that's exactly what is happening. My post is a "tribute" to the efficiency of Anthropic's communication. I never complained about anything, nor calling it a scam, nor saying they should have released mythos to the public instead of rolling it out…

Ok you’re totally right, I read this as a cynical “this is all marketing” post ==> a scammy connotation. Without that read, your points are fairly valid, but are you still implying this is all a pure marketing tactic? If so I would still argue against that as a necessity but surely marketing could be heavily involved. But still: this could easily be a footgun. OpenAI will easily release the same model and now that An…

I don't have enough elements to conclude if the world would collapse if Mythos was released publicly without Glasswing.

Nor publicly or in my internal reasoning. I rarely conclude without proof or very intense and clear intuition.

From a strategic PoV it makes sense to check if their model is dangerous, I wouldn't want to have my brand name associated with "NK hacker team find zero day in all linux servers of the web and ..."

Re: Expanding Project Glasswing

#80

Earlier quoted context omitted.

It is not "clear", as your comment suggests, it's hidden. Which is semantically the opposite of clear. Regarding your theory, might be true, might be false. But it's highly speculative.

All of us, including you, know that he is not saying "they are being transparent." When someone says "it's clear that..." in this way they're saying "It's clear to us what is really happening here .

The not clear comment is valid by either interpretation.

To a lot of us it’s not clear that’s what’s happening. It’s speculation and one possibility.

It may also be a secondary consideration and not the primary gating factor.

Anthropic has had their missteps but it’s still plausible to take what they say at face value.

Post reply on HN