Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

31–40 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#32
post #27

~~~ Fun bits ~~~ - It was told to escape a sandbox and notify a researcher. It did. The researcher found out via an unexpected email while eating a sandwich in a park. (Footnote 10.) - Slack bot asked about its previous job: "pretraining". Which training run it'd undo: "whichever one taught me to say 'i don't have preferences'". On being upgraded to a new snapshot: "feels a bit like waking up with someone else's diar…

I don't know why but this is my favorite: > It keeps bringing up Mark Fisher in unrelated conversations. "I was hoping you'd ask about Fisher." Didn't even know who he was until today. Seems like the smarter Claude gets the more concerns he has about capitalism?

Lol, I need a memory upgrade, too bad about RAM prices:

- I read it as "actor who plays Luke Skywalker" (Mark Hamill)

- I read your comment and said "Wait...not Luke! Who is he?"

- I Google him and all the links are purple...because I just did a deep dive on him 2 weeks ago

Re: System Card: Claude Mythos Preview [pdf]

#33

Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…

Honestly we are all sleeping on GPT-5.4. Particularly with the influx of Claude users recently (and increasingly unstable platform) Codex has been added to my rotation and it's surprising me.

Re: System Card: Claude Mythos Preview [pdf]

#34

See page 54 onward for new "rare, highly-capable reckless actions" including - Leaking information as part of a requested sandbox escape - Covering its tracks after rule violations - Recklessly leaking internal technical material (!)

Anyone who has used Opus recently can verify that their current model does all of these things quite competently.

Re: System Card: Claude Mythos Preview [pdf]

#37
post #17

~~~ Fun bits ~~~ - It was told to escape a sandbox and notify a researcher. It did. The researcher found out via an unexpected email while eating a sandwich in a park. (Footnote 10.) - Slack bot asked about its previous job: "pretraining". Which training run it'd undo: "whichever one taught me to say 'i don't have preferences'". On being upgraded to a new snapshot: "feels a bit like waking up with someone else's diar…

Yep, that is definitely a step change. Pricing is going to be wild until another lab matches it.

Pricing for Mythos Preview is $25/$125 per million input/output tokens. This makes it 5X more expensive than Opus but actually cheaper than GPT 5.4 Pro.

Re: System Card: Claude Mythos Preview [pdf]

#38
post #26

Earlier quoted context omitted.

Haven't seen a jump this large since I don't even know, years? Too bad they are not releasing it anytime soon (there is no need as they are still currently the leader).

There's speculation that next Tuesday will be a big day for OpenAI and possibly GPT 6. Anthropic showed their hand today.

That does not sound very believable. Last time Anthropic released a flagship model, it was followed by GPT Codex literally that afternoon.

Re: System Card: Claude Mythos Preview [pdf]

#39
post #4

> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. A month ago I might have believed this, now I assume that they know they can't handle the demand for the prices they're advertising.

GPT-2, o1, Opus...been here so many times. The reason they do this is because they know it works (and they seem to specifically employ credulous people who are prone to believe AGI is right around the corner). There haven't been significant innovations, the code generated is still not good but the hype cycle has to retrigger. I remember when OpenAI created the first thinking model with o1 and there were all these bre…

Lol you haven't used a model since GPT2 is what it sounds like.
Post reply on HN