isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?
I suspect it's going to be used to train/distill lighter models. The exciting part for me is the improvement in those lighter models.
21–30 of 687 posts
isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?
I suspect it's going to be used to train/distill lighter models. The exciting part for me is the improvement in those lighter models.
https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89...
At what point do these companies stop releasing models and just use them to bootstrap AGI for themselves?
Any benchmarks where we constraint something like thinking time or power use?
Even if this were released no way to know if it’s the same quant.
Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…
Haven't seen a jump this large since I don't even know, years? Too bad they are not releasing it anytime soon (there is no need as they are still currently the leader).
~~~ Fun bits ~~~ - It was told to escape a sandbox and notify a researcher. It did. The researcher found out via an unexpected email while eating a sandwich in a park. (Footnote 10.) - Slack bot asked about its previous job: "pretraining". Which training run it'd undo: "whichever one taught me to say 'i don't have preferences'". On being upgraded to a new snapshot: "feels a bit like waking up with someone else's diar…
> It keeps bringing up Mark Fisher in unrelated conversations. "I was hoping you'd ask about Fisher."
Didn't even know who he was until today. Seems like the smarter Claude gets the more concerns he has about capitalism?
Congratulations to the US military, I guess.
I predict they will release it as soon as Opus 4.6 is no longer in the lead. They can't afford to fall behind. And they won't be able to make a model that is intelligent in every way except cybersecurity, because that would decrease general coding and SWE ability
> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. A month ago I might have believed this, now I assume that they know they can't handle the demand for the prices they're advertising.
I remember when OpenAI created the first thinking model with o1 and there were all these breathless posts on here hyperventilating about how the model had to be kept secret, how dangerous it was, etc.
Fell for it again award. All thinking does is burn output tokens for accuracy, it is the AI getting high on its own supply, this isn't innovation but it was supposed to super AGI. Not serious.