System Card: Claude Mythos Preview [pdf]
91–100 of 687 posts
Re: System Card: Claude Mythos Preview [pdf]
#92Their best model to date and they won’t let the general public use it. This is the first moment where the whole “permanent underclass” meme starts to come into view. I had through previously that we the consumers would be reaping the benefits of these frontier models and now they’ve finally come out and just said it - the haves can access our best, and have-nots will just have use the not-quite-best. Perhaps I was be…
If AI really is bench marking this well -> just sell it as a complete replacement which you can charge for some insane premium, just has to cost less than the employees...
I was worried before, but this is truly the darkest timeline if this is really what these companies are going for.
Re: System Card: Claude Mythos Preview [pdf]
#93Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…
The real part is SWE-bench Verified since there is no way to overfit. That's the only one we can believe.
OpenAI had a whole post about this, where they recommended switching to SWE-bench Pro as a better (but still imperfect) benchmark:
https://openai.com/index/why-we-no-longer-evaluate-swe-bench...
> We audited a 27.6% subset of the dataset that models often failed to solve and found that at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions
> SWE-bench problems are sourced from open-source repositories many model providers use for training purposes. In our analysis we found that all frontier models we tested were able to reproduce the original, human-written bug fix
> improvements on SWE-bench Verified no longer reflect meaningful improvements in models’ real-world software development abilities. Instead, they increasingly reflect how much the model was exposed to the benchmark at training time
> We’re building new, uncontaminated evaluations to better track coding capabilities, and we think this is an important area to focus on for the wider research community. Until we have those, OpenAI recommends reporting results for SWE-bench Pro.
Re: System Card: Claude Mythos Preview [pdf]
#94Earlier quoted context omitted.
A jump that we will never be able to use since we're not part of the seemingly minimum 100 billion dollar company club as requirement to be allowed to use it. I get the security aspect, but if we've hit that point any reasonably sophisticated model past this point will be able to do the damage they claim it can do. They might as well be telling us they're closing up shop for consumer models. They should just say they…
More than killer AI I'm afraid of Anthropic/OpenAI going into full rent-seeking mode so that everyone working in tech is forced to fork out loads of money just to stay competitive on the market. These companies can also choose to give exclusive access to hand picked individuals and cut everyone else off and there would be nothing to stop them. This is already happening to some degree, GPT 5.3 Codex's security capabil…
Re: System Card: Claude Mythos Preview [pdf]
#95> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. A month ago I might have believed this, now I assume that they know they can't handle the demand for the prices they're advertising.
you would be a fool to believe it at any point in time. Amodei is anthropomorphic grease, even more so than Altman. Anthropic is burning through billions of VC cash. if this model was commercially viable, it would've been released yesterday.
Re: System Card: Claude Mythos Preview [pdf]
#96Are you guys ready for the bifurcation when the top models are prohibitively expensive to normal users? If your AI budget $2000+ a month? Or are you going to be part of the permanent free tier underclass?
If one is to believe the API prices are reasonable representation of non subsidized "real world pricing" (with model training being the big exception), then the models are getting cheaper over time. GPT 4.5 was $150.00 / 1M tokens IIRC. GPT o1-pro was $600 / 1M tokens.
Re: System Card: Claude Mythos Preview [pdf]
#97Earlier quoted context omitted.
Anyone who has used Opus recently can verify that their current model does all of these things quite competently.
That has also been my experience. And if Mythos is even worse, unless you have a significantly awesome harness, sounds like pretty unusable if you don't want to risk those problems.
If you look at recent changes in Opus behaviour and this model that is, apparently, amazingly powerful but even more unsafe...seems suspect.
Re: System Card: Claude Mythos Preview [pdf]
#98Ah, so this is how the source code got leaked.
/s
Re: System Card: Claude Mythos Preview [pdf]
#99Earlier quoted context omitted.
Doesn't Anthropic not have that contract anymore, after all that buzz a month or so ago?
The point of that buzz was to force Anthropic to provide Mythos to the military.
Re: System Card: Claude Mythos Preview [pdf]
#100Earlier quoted context omitted.
Inference for the same results has been dropping 10x year over year[0] [0] https://ziva.sh/blogs/llm-pricing-decline-analysis
Sure, but "the same results" will rapidly become unacceptable results if much better results are available.