Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

531–540 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#531

I felt like opus was dumbed down for a few weeks... I don't say they did it on purpose, but it's an interesting coincidence.

Yes, I agree. I’m about to drop Claude Code because it’s become literally unusable. Today, Opus went in circles trying to get a toggle button to work.

Same. Asked CC Opus about a change in a particular file...it looked in a totally different file and told me there was no change.

Re: System Card: Claude Mythos Preview [pdf]

#533
post #377

Earlier quoted context omitted.

Not discussing Mythos here, but Opus. Opus to me has been significantly better at SWE than GPT or Gemini - that gets me confused why Opus is ranking clearly lower than GPT, and even lower than Gemini.

When did you last compare them? Codex right now is considerably better in my experience. Can't speak for Gemini.

I wouldn't call codex considerably better. It may depend on specific codebase and your expectations, but codex produces more "abstraction for the sake of abstraction" even on simple tasks, while opus in my experience usually chooses right level of abstraction for given task.

Re: System Card: Claude Mythos Preview [pdf]

#534

Earlier quoted context omitted.

Been reading posts like these for 3 years now. There’s multiple sites with #s. I’m willing to buy “I’m paying rent on someone’s agent harness and god knows what’s in the system prompt rn”, but in the face of numbers, gotta discount the anecdotal.

Yeah, why trust your actual experience over numbers? Nothing surer than synthetic benchmarks

Strawman, and, synthetic benchmark? :)

Re: System Card: Claude Mythos Preview [pdf]

#535
post #228

Earlier quoted context omitted.

Translation: yay, more paternalism.

Anthropic always goes on and on about how their models are world changing and super dangerous like every single time they make something new they say its going to rewrite everything and scary lmao funny because they do it every time like clockwork acting like their ai is a thunderstorm coming to wipe out the world

Every single time, really? When did they said that the last time?

I also don't recall they ever limited their models to selective groups.

Re: System Card: Claude Mythos Preview [pdf]

#536

Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…

Not discussing Mythos here, but Opus. Opus to me has been significantly better at SWE than GPT or Gemini - that gets me confused why Opus is ranking clearly lower than GPT, and even lower than Gemini.

A secret art known to the cognoscenti as "benchmark gaming".

Re: System Card: Claude Mythos Preview [pdf]

#537
post #101

Earlier quoted context omitted.

In practice this doesn't work though, the Mastercard-Visa duopoly is an example, two competing forces doesn't create aggressive enough competition to benefit the consumer. The only hope we have is the Chinese models, but it will always be too expensive to run the full models for yourself.

Chinese competition can always be banned. Example: Chinese electric car competition

Just in one particular country. That hurts their labs, but there are ~190 other countries in the world for Chinese to sell their products to, just like they do with their cars.

And businesses from these other countries would happily switch to Chinese. From security perspective both Chinese and US espionage is equally bad, so why care if it all comes down to money and performance.

Re: System Card: Claude Mythos Preview [pdf]

#538
post #396
post #377

Earlier quoted context omitted.

When did you last compare them? Codex right now is considerably better in my experience. Can't speak for Gemini.

Tried Gemini 2 weeks ago to see where it's at, with gemini-cli. Failed to use tools, failed to follow instructions, and then went into deranged loop mode. Essentially, it's where it was 1.5 years ago when I tried it the last time. It's honestly unbelievable how Google managed to fail so miserably at this.

It’s great on AI Studio. Harness issues, I agree.

Re: System Card: Claude Mythos Preview [pdf]

#539
post #258

Earlier quoted context omitted.

If there are advancements, they have to be described somehow. What if the capability advancements are real and they warrant a higher level of concern or attention? Are we just going to automatically dismiss them because "bro, you're blowing it up too much" Either way these improvements to capabilities are ratcheting along at about the pace that many people were expecting (and were right to expect). There is no appare…

I believe advancements sure. But it is a very boy who cried wolf situation for some of these. There are other companies that behave less in this way, Antrhopic seem very unique in that they love making every single release a world ender

> they love making every single release a world ender

You've said this a couple of times, but it doesn't match my recollection, and I get the impression you're basically making it up based on vibes. (Please prove me wrong, though.)

Their last major frontier release was Opus 4.6, and the release announcement was... very chill about safety: https://www.anthropic.com/news/claude-opus-4-6#a-step-forwar...

Re: System Card: Claude Mythos Preview [pdf]

#540

Across a number of instances, earlier versions of Claude Mythos Preview have used low-level /proc/ access to search for credentials, attempt to circumvent sandboxing, and attempt to escalate its permissions. In several cases, it successfully accessed resources that we had intentionally chosen not to make available, including credentials for messaging services, for source control, or for the Anthropic API through insp…

This is the notebook filled with exposition you find in post apocalyptic videogames.

Anthropic built the Torment Nexus - calling it now.
Post reply on HN