Uh... what? Does anyone have any idea what these guys are talking about?
System Card: Claude Mythos Preview [pdf]
181–190 of 687 posts
Re: System Card: Claude Mythos Preview [pdf]
#182In French a "mytho" is a mythomaniac. Quite fitting.
Re: System Card: Claude Mythos Preview [pdf]
#183Re: System Card: Claude Mythos Preview [pdf]
#184Interesting reading. They are still focusing on "catastrophic risks" related to chemical and biological weapons production; or misaligned models wreaking havoc. But they are not addressing the elephant in the room: * Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement.
Re: System Card: Claude Mythos Preview [pdf]
#185Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…
but how does it perform on pelican riding a bicycle bench? why are they hiding the truth?! (edit: I hope this is an obvious joke. less facetiously these are pretty jaw dropping numbers)
Re: System Card: Claude Mythos Preview [pdf]
#186Earlier quoted context omitted.
uhh the model found actual vulnerabilities in software that people use. either you believe that the vulnerabilities were not found or were not serious enough to warrant a more thoughtful release
So did GPT-4. https://arxiv.org/html/2402.06664v1 Like think carefully about this. Did they discover AGI? Or did a bunch of investors make a leveraged bet on them "discovering AGI" so they're doing absolutely anything they can to make it seem like this time it's brand new and different. If we're to believe Anthropic on these claims, we also have to just take it on faith, with absolutely no evidence, that they've made…
Also they just hit a $30B run-rate, I don't think they're that needy for new hype cycles.
Re: System Card: Claude Mythos Preview [pdf]
#187Interesting reading. They are still focusing on "catastrophic risks" related to chemical and biological weapons production; or misaligned models wreaking havoc. But they are not addressing the elephant in the room: * Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement.
I think we're pretty good at that without AI.
Re: System Card: Claude Mythos Preview [pdf]
#188Earlier quoted context omitted.
In practice this doesn't work though, the Mastercard-Visa duopoly is an example, two competing forces doesn't create aggressive enough competition to benefit the consumer. The only hope we have is the Chinese models, but it will always be too expensive to run the full models for yourself.
Chinese competition can always be banned. Example: Chinese electric car competition
Re: System Card: Claude Mythos Preview [pdf]
#189> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. A month ago I might have believed this, now I assume that they know they can't handle the demand for the prices they're advertising.
That's for the investors basically. Scarcity and FOMO.