Live data from Hacker News

Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

claude.com

51–56 of 56 posts

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#51

Earlier quoted context omitted.

You didn't really address my point. You just... stated your own.

I definitely did. I was talking about how in a world where open models exist and are very much capable, Anthropic is sort of a joke. You asked if that is true, and speculated about mythos. Uninterested in speculating about mythos, I explained why it is irrelevant anyway. This was a successul conversation.

Glad you think so.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#52
post #40

Earlier quoted context omitted.

As basically every subject matter expert has stated again and again, benchmarks do not tell reveal anything meaningful and have effectively no relation to the model's actual capabilities.

Which is why I actually used all those models. At work I am stuck with Claude because it is what the employer provides. In my home setup I use a combination of different models; mostly GLM-5.3, DS-V4-Flash, and MiMo-2.5-pro. I vastly prefer my home setup.

That says more about your tasks than anything else.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#53

Earlier quoted context omitted.

I’ve been working on a decompilation project that fable was choking on and GLM 5.3 has been chunking away at it for 72 hours now? I think it’s my favorite agentic/implementer model right now.

My big fear is that they are busy nerfing the weights for "safety" before releasing them. In fact, they've more-or-less said as much. I have a feeling what we are about to see on HuggingFace is not the GLM 5.3 that you're using now.

You mean when they said they are doing safety evaluation and hardening? I'm having the same fear as you.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#54
post #52

Earlier quoted context omitted.

Which is why I actually used all those models. At work I am stuck with Claude because it is what the employer provides. In my home setup I use a combination of different models; mostly GLM-5.3, DS-V4-Flash, and MiMo-2.5-pro. I vastly prefer my home setup.

That says more about your tasks than anything else.

You're free to speculate and pretend you had some insight.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#55
post #53

Earlier quoted context omitted.

My big fear is that they are busy nerfing the weights for "safety" before releasing them. In fact, they've more-or-less said as much. I have a feeling what we are about to see on HuggingFace is not the GLM 5.3 that you're using now.

You mean when they said they are doing safety evaluation and hardening? I'm having the same fear as you.

Yes, exactly.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#56

The problem is that "cybersecurity" isn't some special task that only your security team does. In the project I maintain, I find bugs and fix bugs. Some of those bugs might result in an LPE. I generate a regression test, then I fix the bug. The problem is that generating a regression test for that type of bug is technically a PoC. I can almost never get Fable to create one. Sometimes Opus 5 punts as well. Same with S…

Sometimes you can't get Fable to create one? Most of the time I can't even get Fable to investigate why a unit test is crashing because a segfault is a cyber security risk so it changes the model automatically. Drives me hp the wall how hard they clutch their pearls here, Codex has given me no such trouble.
Post reply on HN