Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

61–70 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#61

Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…

The real part is SWE-bench Verified since there is no way to overfit. That's the only one we can believe.

Re: System Card: Claude Mythos Preview [pdf]

#65
post #37
post #17

Earlier quoted context omitted.

Yep, that is definitely a step change. Pricing is going to be wild until another lab matches it.

Pricing for Mythos Preview is $25/$125 per million input/output tokens. This makes it 5X more expensive than Opus but actually cheaper than GPT 5.4 Pro.

Important to note it's only for participants, not the general public.

Re: System Card: Claude Mythos Preview [pdf]

#66
Interesting reading.

They are still focusing on "catastrophic risks" related to chemical and biological weapons production; or misaligned models wreaking havoc.

But they are not addressing the elephant in the room:

* Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement.

Re: System Card: Claude Mythos Preview [pdf]

#67

See page 54 onward for new "rare, highly-capable reckless actions" including - Leaking information as part of a requested sandbox escape - Covering its tracks after rule violations - Recklessly leaking internal technical material (!)

Anyone who has used Opus recently can verify that their current model does all of these things quite competently.

That has also been my experience. And if Mythos is even worse, unless you have a significantly awesome harness, sounds like pretty unusable if you don't want to risk those problems.

Re: System Card: Claude Mythos Preview [pdf]

#68

Are you guys ready for the bifurcation when the top models are prohibitively expensive to normal users? If your AI budget $2000+ a month? Or are you going to be part of the permanent free tier underclass?

Inference for the same results has been dropping 10x year over year[0]

[0] https://ziva.sh/blogs/llm-pricing-decline-analysis

Post reply on HN