Earlier quoted context omitted.
More than killer AI I'm afraid of Anthropic/OpenAI going into full rent-seeking mode so that everyone working in tech is forced to fork out loads of money just to stay competitive on the market. These companies can also choose to give exclusive access to hand picked individuals and cut everyone else off and there would be nothing to stop them. This is already happening to some degree, GPT 5.3 Codex's security capabil…
Well don’t forget we still have competition. Were anthropic to rent seek OpenAI would undercut them. Were OpenAI and anthropic to collude that would be illegal. For anthropic to capture the entire coding agent market and THEN rent seek, these days it’s never been easier to raise $1B and start a competing lab
System Card: Claude Mythos Preview [pdf]
101–110 of 687 posts
Re: System Card: Claude Mythos Preview [pdf]
#102Re: System Card: Claude Mythos Preview [pdf]
#103Earlier quoted context omitted.
Lol you haven't used a model since GPT2 is what it sounds like.
Just checked my subscription start date for Anthropic. September 2023, I believe before they announced public launch. Sorry kid.
You must see some value, or are you in a situation where you're required to test / use it, eg to report on it or required by employer?
(I would disagree about the code, the benefits seem obvious to me. But I'm still curious why others would disagree, especially after actively using them for years.)
Re: System Card: Claude Mythos Preview [pdf]
#104Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…
We're gonna need some new benchmarks... ARC-AGI-3 might be the only remaining benchmark below 50%
Here is an example question: https://i.redd.it/5jl000p9csee1.jpeg
No human could even score 5% on HLE.
Re: System Card: Claude Mythos Preview [pdf]
#105Earlier quoted context omitted.
GPT is shit at writing code. It's not dumb - extra high thinking is really good at catching stuff - but it's like letting a smart junior into your codebase - ignore all the conventions, surrounding context, just slop all over the place to get it working. Claude is just a level above in terms of editing code.
Yes, it's becoming clear that OpenAI kinda sucks at alignment. GPT-5 can pass all the benchmarks but it just doesn't "feel good" like Claude or Gemini.
Re: System Card: Claude Mythos Preview [pdf]
#106Interesting reading. They are still focusing on "catastrophic risks" related to chemical and biological weapons production; or misaligned models wreaking havoc. But they are not addressing the elephant in the room: * Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement.
Re: System Card: Claude Mythos Preview [pdf]
#107isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?
I don’t doubt they have found interesting security holes, the question is how they actually found them.
This System Card is just a sales whitepaper and just confirms what that “leak” from a week or so ago implied.
Re: System Card: Claude Mythos Preview [pdf]
#108At what point do these companies stop releasing models and just use them to bootstrap AGI for themselves?
Re: System Card: Claude Mythos Preview [pdf]
#109Earlier quoted context omitted.
So... you're not excited because it might take a few months before we can use it or something? I don't get your comment.
I think the general question is if they'll release it at all, haven't yet read anything stating that they would
https://en.wikipedia.org/wiki/Capitalism
https://en.wikipedia.org/wiki/Race_to_the_bottom
https://en.wikipedia.org/wiki/Arms_race
Of course they'll release it once they can de-risk it sufficently and/or a competitor gets close enough on their tail, whichever comes first.
Re: System Card: Claude Mythos Preview [pdf]
#110> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. Absolutely genius move from Anthropic here. This is clearly their GPT-4.5, probably 5x+ the size of their best current models and way too expensive to subsidize on a subscription for only marginal gains in real world scenarios. But unlike OpenAI, they have the level of hysteric marketing hype required to…
From Stratechery[0]:
> Strategy Credit: An uncomplicated decision that makes a company look good relative to other companies who face much more significant trade-offs. For example, Android being open source