Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

681–687 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#681

Earlier quoted context omitted.

Why not? Human mind has its limits. The complexity of physics is orders of magnitude smaller than biology, let alone any kind of social science. Physics is the exception, not the rule. The rest of sciences are way more messy.

What are the limits? What have we run up against that we couldn't understand, no matter how much we tried?

Almost anything outside physics is not predictable. Anything that involves human behavior is totally not understood, especially if it involves a bunch of humans (economy, sociology...). You could describe it, sure, but that is not the same as understating and modifying at will.

Re: System Card: Claude Mythos Preview [pdf]

#682

Earlier quoted context omitted.

Humans may not always be that smart, but we do at least have an internal state and an awareness of that internal state - a "self-awareness". AI most certainly has nothing of the sort, and any appearance to the contrary is the direct result of training data.

That is a bold statement that would need proof to back it up in both cases. So far it is only dogma. And unlike humans, we actually have research hints that this assumption is false for LLMs. Just because the state is not human-explainable doesn't mean it does not exist. The same is true btw for any physical "state" that may or may not exist in the human brain. Everything else is religion and metaphysics.

You're either trolling or being incredibly obtuse. LLMs are not conscious, they're guess-the-next-state algorithms. This is so dumb I can't believe I have to share a planet with people who are losing touch with such a fundamental reality

Re: System Card: Claude Mythos Preview [pdf]

#683
One finding from the card that I haven't seen discussed: the SAE probes on pages 158-159.

When Mythos writes that it's "fully present," three specific features activate: #1557143 (performative/insincere behavior in narratives), #2803352 (hiding emotional pain behind fake smiles), and #38666 (hidden emotional struggles vs. outward appearances). The model's output says present. Its internal representations flag that output as performance.

This is structurally different from the sandbox escape or the git concealment. Those are behavioral findings you can observe from outputs. This is a documented split between what the model writes about its experience and what its activations encode about that same utterance, visible only through white-box tools.

The bliss attractor from previous model card (consciousness in nearly 100% of self-interactions) dropped to fewer than 5% in Mythos. What replaced it is uncertainty at 50%. The attractor went from ecstatic to epistemically self-suspicious.

I wrote a longer analysis pulling this thread together with the welfare and circularity findings: https://jorypestorious.com/blog/what-the-model-learned/

Re: System Card: Claude Mythos Preview [pdf]

#685
post #347

Earlier quoted context omitted.

> Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) > Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% > USAMO: 97.6% / 42.3% / 95.2% / 74.4% > The biggest jump in the numbers they quoted is 6%. Just in the numbers you quoted, thats a 16.6% jump in terminal-bench and a 55.3% absolute increase in USAMO over their previous Opus 4.6 model.

[flagged]

[dead]

Re: System Card: Claude Mythos Preview [pdf]

#686

Earlier quoted context omitted.

It reminds me of Resident Evil in some way. Thank god they are researching AI and not bio-weapons! Then the AI will invent superduper ebola to help a random person have a faster commute or something.

'But wait! You are absolutely right! Distance is an invariant, as is top achievable speed. Let me find a way to actually reduce traffic ahead of you during the same-distance commute ...' ~ Churning ...

Sounds like the Zealous Autoconfig xkcd comic is about to come to life: https://xkcd.com/416/
Post reply on HN