I felt like opus was dumbed down for a few weeks... I don't say they did it on purpose, but it's an interesting coincidence.
Yes, I agree. I’m about to drop Claude Code because it’s become literally unusable. Today, Opus went in circles trying to get a toggle button to work.
System Card: Claude Mythos Preview [pdf]
531–540 of 687 posts
Re: System Card: Claude Mythos Preview [pdf]
#532[flagged]
Re: System Card: Claude Mythos Preview [pdf]
#533Earlier quoted context omitted.
Not discussing Mythos here, but Opus. Opus to me has been significantly better at SWE than GPT or Gemini - that gets me confused why Opus is ranking clearly lower than GPT, and even lower than Gemini.
When did you last compare them? Codex right now is considerably better in my experience. Can't speak for Gemini.
Re: System Card: Claude Mythos Preview [pdf]
#534Earlier quoted context omitted.
Been reading posts like these for 3 years now. There’s multiple sites with #s. I’m willing to buy “I’m paying rent on someone’s agent harness and god knows what’s in the system prompt rn”, but in the face of numbers, gotta discount the anecdotal.
Yeah, why trust your actual experience over numbers? Nothing surer than synthetic benchmarks
Re: System Card: Claude Mythos Preview [pdf]
#535Earlier quoted context omitted.
Translation: yay, more paternalism.
Anthropic always goes on and on about how their models are world changing and super dangerous like every single time they make something new they say its going to rewrite everything and scary lmao funny because they do it every time like clockwork acting like their ai is a thunderstorm coming to wipe out the world
I also don't recall they ever limited their models to selective groups.
Re: System Card: Claude Mythos Preview [pdf]
#536Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…
Not discussing Mythos here, but Opus. Opus to me has been significantly better at SWE than GPT or Gemini - that gets me confused why Opus is ranking clearly lower than GPT, and even lower than Gemini.
Re: System Card: Claude Mythos Preview [pdf]
#537Earlier quoted context omitted.
In practice this doesn't work though, the Mastercard-Visa duopoly is an example, two competing forces doesn't create aggressive enough competition to benefit the consumer. The only hope we have is the Chinese models, but it will always be too expensive to run the full models for yourself.
Chinese competition can always be banned. Example: Chinese electric car competition
And businesses from these other countries would happily switch to Chinese. From security perspective both Chinese and US espionage is equally bad, so why care if it all comes down to money and performance.
Re: System Card: Claude Mythos Preview [pdf]
#538Earlier quoted context omitted.
When did you last compare them? Codex right now is considerably better in my experience. Can't speak for Gemini.
Tried Gemini 2 weeks ago to see where it's at, with gemini-cli. Failed to use tools, failed to follow instructions, and then went into deranged loop mode. Essentially, it's where it was 1.5 years ago when I tried it the last time. It's honestly unbelievable how Google managed to fail so miserably at this.
Re: System Card: Claude Mythos Preview [pdf]
#539Earlier quoted context omitted.
If there are advancements, they have to be described somehow. What if the capability advancements are real and they warrant a higher level of concern or attention? Are we just going to automatically dismiss them because "bro, you're blowing it up too much" Either way these improvements to capabilities are ratcheting along at about the pace that many people were expecting (and were right to expect). There is no appare…
I believe advancements sure. But it is a very boy who cried wolf situation for some of these. There are other companies that behave less in this way, Antrhopic seem very unique in that they love making every single release a world ender
You've said this a couple of times, but it doesn't match my recollection, and I get the impression you're basically making it up based on vibes. (Please prove me wrong, though.)
Their last major frontier release was Opus 4.6, and the release announcement was... very chill about safety: https://www.anthropic.com/news/claude-opus-4-6#a-step-forwar...
Re: System Card: Claude Mythos Preview [pdf]
#540Across a number of instances, earlier versions of Claude Mythos Preview have used low-level /proc/ access to search for credentials, attempt to circumvent sandboxing, and attempt to escalate its permissions. In several cases, it successfully accessed resources that we had intentionally chosen not to make available, including credentials for messaging services, for source control, or for the Anthropic API through insp…
This is the notebook filled with exposition you find in post apocalyptic videogames.