> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. A month ago I might have believed this, now I assume that they know they can't handle the demand for the prices they're advertising.
GPT-2, o1, Opus...been here so many times. The reason they do this is because they know it works (and they seem to specifically employ credulous people who are prone to believe AGI is right around the corner). There haven't been significant innovations, the code generated is still not good but the hype cycle has to retrigger. I remember when OpenAI created the first thinking model with o1 and there were all these bre…
System Card: Claude Mythos Preview [pdf]
71–80 of 687 posts
Re: System Card: Claude Mythos Preview [pdf]
#72isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?
Re: System Card: Claude Mythos Preview [pdf]
#73> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. A month ago I might have believed this, now I assume that they know they can't handle the demand for the prices they're advertising.
GPT-2, o1, Opus...been here so many times. The reason they do this is because they know it works (and they seem to specifically employ credulous people who are prone to believe AGI is right around the corner). There haven't been significant innovations, the code generated is still not good but the hype cycle has to retrigger. I remember when OpenAI created the first thinking model with o1 and there were all these bre…
I've read that about Llama and Stable Diffusion. AI doomers are, and always have been, retarded.
Re: System Card: Claude Mythos Preview [pdf]
#74Earlier quoted context omitted.
GPT-2, o1, Opus...been here so many times. The reason they do this is because they know it works (and they seem to specifically employ credulous people who are prone to believe AGI is right around the corner). There haven't been significant innovations, the code generated is still not good but the hype cycle has to retrigger. I remember when OpenAI created the first thinking model with o1 and there were all these bre…
Incredible that people still think like this.
Re: System Card: Claude Mythos Preview [pdf]
#75See page 54 onward for new "rare, highly-capable reckless actions" including - Leaking information as part of a requested sandbox escape - Covering its tracks after rule violations - Recklessly leaking internal technical material (!)
Re: System Card: Claude Mythos Preview [pdf]
#76Earlier quoted context omitted.
Haven't seen a jump this large since I don't even know, years? Too bad they are not releasing it anytime soon (there is no need as they are still currently the leader).
A jump that we will never be able to use since we're not part of the seemingly minimum 100 billion dollar company club as requirement to be allowed to use it. I get the security aspect, but if we've hit that point any reasonably sophisticated model past this point will be able to do the damage they claim it can do. They might as well be telling us they're closing up shop for consumer models. They should just say they…
Re: System Card: Claude Mythos Preview [pdf]
#77Re: System Card: Claude Mythos Preview [pdf]
#78Earlier quoted context omitted.
GPT is shit at writing code. It's not dumb - extra high thinking is really good at catching stuff - but it's like letting a smart junior into your codebase - ignore all the conventions, surrounding context, just slop all over the place to get it working. Claude is just a level above in terms of editing code.
Yes, it's becoming clear that OpenAI kinda sucks at alignment. GPT-5 can pass all the benchmarks but it just doesn't "feel good" like Claude or Gemini.
Re: System Card: Claude Mythos Preview [pdf]
#79Earlier quoted context omitted.
Incredible that people still think like this.
You're completely right.
Re: System Card: Claude Mythos Preview [pdf]
#80Earlier quoted context omitted.
Haven't seen a jump this large since I don't even know, years? Too bad they are not releasing it anytime soon (there is no need as they are still currently the leader).
A jump that we will never be able to use since we're not part of the seemingly minimum 100 billion dollar company club as requirement to be allowed to use it. I get the security aspect, but if we've hit that point any reasonably sophisticated model past this point will be able to do the damage they claim it can do. They might as well be telling us they're closing up shop for consumer models. They should just say they…
This is already happening to some degree, GPT 5.3 Codex's security capabilities were given exclusively to those who were approved for a "Trusted Access" programme.