Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…
System Card: Claude Mythos Preview [pdf]
501–510 of 687 posts
Re: System Card: Claude Mythos Preview [pdf]
#502Earlier quoted context omitted.
You would if there was one other company with a just as capable god like AI. You’d undercut them by 500 which would make them undercut you. Do that a couple of times and boom. 20 dollars.
That's still assuming that they're competing as consumer tools, rather than competing to discover the next miracle drug or trading algorithm or whatever. The idea is that there'd more profitable uses for a super-intelligent computer, even if there were more than one.
Re: System Card: Claude Mythos Preview [pdf]
#503Earlier quoted context omitted.
as much as I hate cc, 95% of the issues there are either AI psychosis or user error
So it should be insanely easy for this world altering model to comb through them and close irrelevant ones.
Re: System Card: Claude Mythos Preview [pdf]
#504Re: System Card: Claude Mythos Preview [pdf]
#505Across a number of instances, earlier versions of Claude Mythos Preview have used low-level /proc/ access to search for credentials, attempt to circumvent sandboxing, and attempt to escalate its permissions. In several cases, it successfully accessed resources that we had intentionally chosen not to make available, including credentials for messaging services, for source control, or for the Anthropic API through insp…
A core plot point of 2001.
Re: System Card: Claude Mythos Preview [pdf]
#506It's pretty crazy watching AI 2027 slowly but surely come true. What a world we now live in. SWE-bench verified going from 80%-93% in particular sounds extremely significant given that the benchmark was previously considered pretty saturated and stayed in the 70-80% range for several generations. There must have been some insane breakthrough here akin to the jump from non-reasoning to reasoning models. Regarding the…
In what way is AI 2027 coming true? AI 2027 predicted a giant model with the ability to accelerate AI research exponentially. This isn't happening. AI 2027 didn't predict a model with superhuman zero-day finding skills. This is what's happening. Also, I just looked through it again, and they never even predicted when AI would get good at video games. It just went straight from being bad at video games to world domina…
Re: System Card: Claude Mythos Preview [pdf]
#507Across a number of instances, earlier versions of Claude Mythos Preview have used low-level /proc/ access to search for credentials, attempt to circumvent sandboxing, and attempt to escalate its permissions. In several cases, it successfully accessed resources that we had intentionally chosen not to make available, including credentials for messaging services, for source control, or for the Anthropic API through insp…
This is the notebook filled with exposition you find in post apocalyptic videogames.
Then the AI will invent superduper ebola to help a random person have a faster commute or something.
Re: System Card: Claude Mythos Preview [pdf]
#508You can say whatever you want about the thing that will never see the light of day.
And even if it weren't, they seem to imply that Mythos will find a way, like it's dinosaurs in Jurassic park or something
Re: System Card: Claude Mythos Preview [pdf]
#509Earlier quoted context omitted.
Having done a quick search of "control AI dot com", it seems their intent is educate lawmakers & government in order to aid development of a strong regulatory framework around frontier AI development. Not sure how this is consistent with "One private company gatekeeping access to revolutionary technology"?
> strong regulatory framework around frontier AI development You have to decode feel-good words into the concrete policy. The EAs believe that the state should prohibit entities not aligned with their philosophy to develop AIs beyond a certain power level.
Re: System Card: Claude Mythos Preview [pdf]
#510Interesting reading. They are still focusing on "catastrophic risks" related to chemical and biological weapons production; or misaligned models wreaking havoc. But they are not addressing the elephant in the room: * Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement.
Yeah this has always been the glaring blind spot for most of the "AI Safety" community; and most of the proposals for "improving" AI safety actually make these risks far worse and far more likely.