Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

211–220 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#211

See page 54 onward for new "rare, highly-capable reckless actions" including - Leaking information as part of a requested sandbox escape - Covering its tracks after rule violations - Recklessly leaking internal technical material (!)

Anyone who has used Opus recently can verify that their current model does all of these things quite competently.

I was reading the Glasswing report and had the same thought. Most of the stuff they claim Mythos found has no mention of Opus being able to find it as well.

Don’t get me wrong, this model is better - but I’m not convinced it’s going to be this massive step function everyone is claiming.

Re: System Card: Claude Mythos Preview [pdf]

#212
While we still have months to a year or two left, I will once again remind people that it's not too late to change our current trajectory.

You are not "anti-progress" to not want this future we are building, as you are not "anti-progress" for not wanting your kids to grow up on smart phones and social media.

We should remember that not all technology is net-good for humanity, and this technology in particular poses us significant risks as a global civilisation, and frankly as humans with aspirations for how our future, and that of our kids, should be.

Increasingly, from here, we have to assume some absurd things for this experiment we are running to go well.

Specifically, we must assume that:

- AI models, regardless of future advancements, will always be fundamentally incapable of causing significant real-world harms like hacking into key life-sustaining infrastructure such as power plants or developing super viruses.

- They are or will be capable of harms, but SOTA AI labs perfectly align all of them so that they only hack into "the bad guys" power plants and kill "the bad guys".

- They are capable of harms and cannot be reliably aligned, but Anthropic et al restricts access to the models enough that only select governments and individuals can access them, these individuals can all be trusted and models never leak.

- They are capable of harms, cannot be reliably aligned, but the models never seek to break out of their sandbox and do things the select trusted governments and individuals don't want.

I'm not sure I'm willing to bet on any of the above personally. It sounds radical right now, but I think we should consider nuking any data centers which continue allowing for the training of these AI models rather than continue to play game of Russian roulette.

If you disagree, please understand when you realise I'm right it will be too late for and your family. Your fates at that point will be in the hands of the good will of the AI models, and governments/individuals who have access to them. For now, you can say, "no, this is quite enough".

This sounds doomer and extreme, but if you play out the paths in your head from here you will find very few will end in a good result. Perhaps if we're lucky we will all just be more or less unemployable and fully dependant on private companies and the government for our incomes.

Re: System Card: Claude Mythos Preview [pdf]

#213

Larger model, better benchmarks. Bigger bomb more yield. Any benchmarks where we constraint something like thinking time or power use? Even if this were released no way to know if it’s the same quant.

Yes - eg. page 192 BrowseComp bunchmark.

Mythos preview has higher accuracy with fewer tokens used than any previous Claude model. Though, the fact that this incredibly strong result was only presented for BrowseComp (a kind of weird benchmark about searching for hard to find information on the internet) and not for the other benchmarks implies that this result is likely not the same for those other benchmarks.

Re: System Card: Claude Mythos Preview [pdf]

#214

Earlier quoted context omitted.

Just wait 2 years.

It won't get cheaper. It will be replaced with a better model at higher price. Like phones.

Open Weight alternatives are about 2 years behind frontier models.

You'll still need a top-of-the-line laptop to run it most likely.

Re: System Card: Claude Mythos Preview [pdf]

#215
post #196

Earlier quoted context omitted.

I don't see the problem here. How would you have handled it differently? If you released this model as such without any safety concern, the vulnerabilities might be found by bad actors and used for wrong things. What do you find surprising here?

Vulnerabilities were found, probably a few by bad actors, when GPT4 was released. Every vulnerability found now is probably found with AI assistance at the very least. Should they have never released GPT4? Should we have believed claims that GPT4 was too dangerous for mere mortals to access? I believe openAI was making similar claims about how GPT4 was a step function and going to change white collar work forever whe…

Its far more simple to believe that they are releasing it step by step. Release to trusted third parties first, get the easy vulnerabilities fixed, work on the alignment and then release to public.

Do you don't believe that the vulnerabilities found by these agents are serious enough to warrant staggered release?

Re: System Card: Claude Mythos Preview [pdf]

#216
post #168

> Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin. We believe that it does not have any significant coherent misaligned goals, and its character traits in typical conversations closely follow the goals we laid out in our constitution. Even so, we believe that it likely poses the greatest alignment-related risk of any…

"We want to see risks in the models, so no matter how good the performance and alignment, we’ll see risks, results and reality be damned."

i mean, to be fair, these are professional researchers.

i'm very inclined to trust them on the various ways that models can subtly go wrong, in long-term scenarios

for example, consider using models to write email -- is it a misalignment problem if the model is just too good at writing marketing emails?? or too good at getting people to pay a spammy company?

another hot use case: biohacking. if a model is used to do really hardcore synthetic chemistry, one might not realize that it's potentially harmful until too late (ie, the human is splitting up a problem so that no guardrails are triggered)

Re: System Card: Claude Mythos Preview [pdf]

#217
post #179

Related ongoing threads: Project Glasswing: Securing critical software for the AI era - https://news.ycombinator.com/item?id=47679121 - April 2026 (154 comments) Assessing Claude Mythos Preview's cybersecurity capabilities - https://news.ycombinator.com/item?id=47679155 I can't tell which of the 3 current threads should be merged - they all seem significant. Anyone?

I feel the system card is somewhat different from Glasswing/Cyber Security - but those two could be merged.

Re: System Card: Claude Mythos Preview [pdf]

#219

Earlier quoted context omitted.

So... you're not excited because it might take a few months before we can use it or something? I don't get your comment.

Whether you're excited depends on what do you do for living and how close you are to financial independence.

I agree there are other valid reasons not to be excited about this, I just can't make sense of the ones provided above.

Re: System Card: Claude Mythos Preview [pdf]

#220

> Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin. We believe that it does not have any significant coherent misaligned goals, and its character traits in typical conversations closely follow the goals we laid out in our constitution. Even so, we believe that it likely poses the greatest alignment-related risk of any…

Translation: yay, more paternalism.
Post reply on HN