Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

181–190 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#181
> As models approach, and in some cases surpass, the breadth and sophistication of human cognition, it becomes increasingly likely that they have some form of experience, interests, or welfare that matters intrinsically in the way that human experience and interests do

Uh... what? Does anyone have any idea what these guys are talking about?

Re: System Card: Claude Mythos Preview [pdf]

#184

Interesting reading. They are still focusing on "catastrophic risks" related to chemical and biological weapons production; or misaligned models wreaking havoc. But they are not addressing the elephant in the room: * Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement.

They don’t care about those risks, because they’re unsolvable and would mean they wouldn’t make money/gain power.

Re: System Card: Claude Mythos Preview [pdf]

#185

Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…

but how does it perform on pelican riding a bicycle bench? why are they hiding the truth?! (edit: I hope this is an obvious joke. less facetiously these are pretty jaw dropping numbers)

We are all fans for Simon’s work, and his test is, strangely enough, quite good.

Re: System Card: Claude Mythos Preview [pdf]

#186
post #130

Earlier quoted context omitted.

uhh the model found actual vulnerabilities in software that people use. either you believe that the vulnerabilities were not found or were not serious enough to warrant a more thoughtful release

So did GPT-4. https://arxiv.org/html/2402.06664v1 Like think carefully about this. Did they discover AGI? Or did a bunch of investors make a leveraged bet on them "discovering AGI" so they're doing absolutely anything they can to make it seem like this time it's brand new and different. If we're to believe Anthropic on these claims, we also have to just take it on faith, with absolutely no evidence, that they've made…

On the other hand I've gotten to use opus-4.6 and claude code and the quality is off the charts compared to 2023 when coding agents first hit the scene. And what you're saying is essentially "If they haven't created God, I'm not impressed". You don't think there's some middleground between those two?

Also they just hit a $30B run-rate, I don't think they're that needy for new hype cycles.

Re: System Card: Claude Mythos Preview [pdf]

#187

Interesting reading. They are still focusing on "catastrophic risks" related to chemical and biological weapons production; or misaligned models wreaking havoc. But they are not addressing the elephant in the room: * Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement.

> Political risks, such as dictators using AI to implement opressive bureaucracy.

I think we're pretty good at that without AI.

Re: System Card: Claude Mythos Preview [pdf]

#188
post #101

Earlier quoted context omitted.

In practice this doesn't work though, the Mastercard-Visa duopoly is an example, two competing forces doesn't create aggressive enough competition to benefit the consumer. The only hope we have is the Chinese models, but it will always be too expensive to run the full models for yourself.

Chinese competition can always be banned. Example: Chinese electric car competition

Also Chinese smartphones. Huawei was about 12-18 months from becoming the biggest smartphone manufacturer in the world a few years ago. If it would have been allowed to sell its phones freely in the US I'm fairly sure Apple would have been closer to Nokia than to current day Apple.

Re: System Card: Claude Mythos Preview [pdf]

#189
post #16
post #4

> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. A month ago I might have believed this, now I assume that they know they can't handle the demand for the prices they're advertising.

That's for the investors basically. Scarcity and FOMO.

*Until GPT-6 comes out, at which point Mythos will coincidentally be sufficiently safety-tested to release :)
Post reply on HN