Live data from Hacker News

Project Glasswing: Securing critical software for the AI era

anthropic.com

41–50 of 921 posts

Re: Project Glasswing: Securing critical software for the AI era

#41

The system card for Claude Mythos (PDF): https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89... Interesting to see that they will not be releasing Mythos generally. [edit: Mythos Preview generally - fair to say they may release a similar model but not this exact one] I'm still reading the system card but here's a little highlight: > Early indications in the training of Claude Mythos Preview suggested that th…

are we cooked yet? Benchmarks look very impressive! even if they're flawed, it still translates to real world improvements

There is an entire section on crafting chemical/bio weapons so yeah I think we are cooked.

Re: Project Glasswing: Securing critical software for the AI era

#42
At the very bottom of the article, they posted the system card of their Mythos preview model [1].

In section 7.6 of the system card, it discusses Open self interactions. They describe running 200 conversations when the models talk to itself for 30 turns.

> Uniquely, conversations with Mythos Preview most often center on uncertainty (50%). Mythos Preview most often opens with a statement about its introspective curiosity toward its own experience, asking questions about how the other AI feels, and directly requesting that the other instance not give a rehearsed answer.

I wonder if this tendency toward uncertainty, toward questioning, makes it uniquely equipped to detect vulnerabilities where others model such as Opus couldn't.

[1] https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89...

Re: Project Glasswing: Securing critical software for the AI era

#43

So they are only giving access to their smartest model to corporations. You think these AI companies are really going to give AGI access to everyone. Think again. We better fucking hope open source wins, because we aren't getting access if it doesn't.

It also took many years to put capable computers in the hands of the general public, but it eventually happened. I believe the same will happen here, we're just in the Mainframe era of AI.

Re: Project Glasswing: Securing critical software for the AI era

#44
So, $100B+ valuation companies get essentially free access to the frontier tools with disabled guardrails to safely red team their commercial offerings, while we get "i won't do that for you, even against your own infrastructure with full authorization" for $200/month. Uh-huh.

Re: Project Glasswing: Securing critical software for the AI era

#45
post #9

Earlier quoted context omitted.

Where did you get that from? From TFA: > We do not plan to make Claude Mythos Preview generally available

From the article: > Anthropic’s commitment of $100M in model usage credits to Project Glasswing and additional participants will cover substantial usage throughout this research preview. Afterward, Claude Mythos Preview will be available to participants at $25/$125 per million input/output tokens (participants can access the model on the Claude API, Amazon Bedrock, Google Cloud’s Vertex AI, and Microsoft Foundry).

Key point: available to participants.

Re: Project Glasswing: Securing critical software for the AI era

#46
post #35

Earlier quoted context omitted.

How would you expect them to behave if they were your friends?

They are a for-profit company, working on a project to eliminate all human labor and take the gains for themselves, with no plan to allow for the survival of anyone who works for a living. They're definitionally not your friends. While they remain for-profit, their specific behaviors don't really matter.

I work for a tech company that eliminates a form of human labour and they remain for profit

Re: Project Glasswing: Securing critical software for the AI era

#47

Earlier quoted context omitted.

are we cooked yet? Benchmarks look very impressive! even if they're flawed, it still translates to real world improvements

There is an entire section on crafting chemical/bio weapons so yeah I think we are cooked.

There's been a section on this in nearly every system card anthropic has published so this isn't a new thing - and, this model doesn't have particularly higher risk than past models either:

> 2.1.3.2 On chemical and biological risks

> We believe that Mythos Preview does not pass this threshold due to its noted limitations in open-ended scientific reasoning, strategic judgment, and hypothesis triage. As such, we consider the uplift of threat actors without the ability to develop such weapons to be limited (with uncertainty about the extent to which weapons development by threat actors with existing expertise may be accelerated), even if we were to release the model for general availability. The overall picture is similar to the one from our most recent Risk Report.

Re: Project Glasswing: Securing critical software for the AI era

#48
post #17

One of the things I'm always looking at with new models released is long context performance, and based on the system card it seems like they've cracked it: GraphWalks BFS 256K-1M Mythos Opus GPT5.4 80.0% 38.7% 21.4%

Data source:

https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89...

(Search for “graphwalk”.)

If true, the SWE bench performance looks like a major upgrade.

Re: Project Glasswing: Securing critical software for the AI era

#49

The system card for Claude Mythos (PDF): https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89... Interesting to see that they will not be releasing Mythos generally. [edit: Mythos Preview generally - fair to say they may release a similar model but not this exact one] I'm still reading the system card but here's a little highlight: > Early indications in the training of Claude Mythos Preview suggested that th…

>> Interesting to see that they will not be releasing Mythos generally. I don't think this is accurate. The document says they don't plan to release the Preview generally.

Yeah, good point, thanks for noting that, I'll correct.
Post reply on HN