Earlier quoted context omitted.
A jump that we will never be able to use since we're not part of the seemingly minimum 100 billion dollar company club as requirement to be allowed to use it. I get the security aspect, but if we've hit that point any reasonably sophisticated model past this point will be able to do the damage they claim it can do. They might as well be telling us they're closing up shop for consumer models. They should just say they…
I think they already said somewhere that they can't release Mythos because it requires absurdly large amounts of compute. The economics of releasing it just don't work.
System Card: Claude Mythos Preview [pdf]
641–650 of 687 posts
Re: System Card: Claude Mythos Preview [pdf]
#642I wonder what the relationship is between a model's capability and the personality it develops. Page 202: > In interactions with subagents, internal users sometimes observed that Mythos Preview appeared “disrespectful” when assigning tasks. It showed some tendency to use commands that could be read as “shouty” or dismissive, and in some cases appeared to underestimate subagent intelligence by overexplaining trivial t…
Could you transcribe the emoji? HN strips them out.
Re: System Card: Claude Mythos Preview [pdf]
#643Earlier quoted context omitted.
Yes, I agree. I’m about to drop Claude Code because it’s become literally unusable. Today, Opus went in circles trying to get a toggle button to work.
Same. Asked CC Opus about a change in a particular file...it looked in a totally different file and told me there was no change.
It's not necessarily back to where it was, but it's not desk-flipping bad.
Re: System Card: Claude Mythos Preview [pdf]
#644Earlier quoted context omitted.
LLMs have zero metacognition. Don't be fooled - their output is stochastic inference and they have no self-awareness. The best you'll see is an improvised post-hoc rationalization story.
You can turn all these argents around and prove the same is true for humans. Don't be fooled by dogmatic people who spread the idea that the human mind is the pinnacle of cognition in the universe. Best to leave that to religion.
AI most certainly has nothing of the sort, and any appearance to the contrary is the direct result of training data.
Re: System Card: Claude Mythos Preview [pdf]
#645Earlier quoted context omitted.
LLMs have zero metacognition. Don't be fooled - their output is stochastic inference and they have no self-awareness. The best you'll see is an improvised post-hoc rationalization story.
> The best you'll see is an improvised post-hoc rationalization story. Funny, because "post-hoc rationalization" is how many neuroscientists think humans operate. That LLMs are stochastic inference engines is obvious by construction, but you skipped the step where you proved that human thoughts, self-awareness and metacognition are not reducible to stochastic inference.
Re: System Card: Claude Mythos Preview [pdf]
#646Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…
Re: System Card: Claude Mythos Preview [pdf]
#647Earlier quoted context omitted.
https://marginlab.ai/trackers/claude-code-historical-perform...
This just looks like random noise to me? Is it also random on short timespans, like running it 10x in a row?
Re: System Card: Claude Mythos Preview [pdf]
#648Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…
Re: System Card: Claude Mythos Preview [pdf]
#649> Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin. We believe that it does not have any significant coherent misaligned goals, and its character traits in typical conversations closely follow the goals we laid out in our constitution. Even so, we believe that it likely poses the greatest alignment-related risk of any…
Alignment “appearing” better as model capabilities increase scares the shit out of me, tbh.
Re: System Card: Claude Mythos Preview [pdf]
#650I've long maintained that the real indicator that AGI is imminent is that public availability stops being a thing. If you truly believed you had a superhuman, godlike mind in your thrall, renting it out for $20/month would be the last thing you would choose to do with it.
Simpler explanation : they don't have enough GPUs to release this much larger model.