Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

391–400 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#391

Earlier quoted context omitted.

> Given that for a number of these benchmarks, it seems to be barely competitive with the previous gen We're not reading the same numbers I think. Compared to Opus 4.6, it's a big jump nearly in every single bench GP posted. They're "only" catching up to Google's Gemini on GPQA and MMMLU but they're still beating their own Opus 4.6 results on these two. This sounds like a much better model than Opus 4.6.

> We're not reading the same numbers I think. We must not be. That's why I listed out the ones where it is barely competitive from @babelfish's table, which itself is extracted from Pg 186 & 187 of the System Card, which has the comparison with Opus 4.6, GPT 5.4 and Gemini 3.1 Pro. Sure, it may be better than Opus 4.6 on some of those, but barely achieves a small increase over GPT-5.4 on the ones I called out.

> barely competitive

It's higher than all other models except vs Gemini 3.1 Pro on MMMLU

MMMLU is generally thought to be maxed out - as it it might not be possible to score higher than those scores.

> Overall, they estimated that 6.5% of questions in MMLU contained an error, suggesting the maximum attainable score was significantly below 100%[1]

Other models get close on GPQA Diamond, but it wouldn't be surprising to anyone if the max possible on that was around the 95% the top models are scoring.

[1] https://en.wikipedia.org/wiki/MMLU

Re: System Card: Claude Mythos Preview [pdf]

#392
post #208

Earlier quoted context omitted.

Alignment “appearing” better as model capabilities increase scares the shit out of me, tbh.

Conversely: in humans, intelligence is inversely correlated with crime. It doesn't go to zero, however!

Is that actually well defined given the very low sample size at the top?

To the best of my knowledge, none of the individuals believed to have an IQ >200 have committed an actual crime.

The closest I found is William James Sidis's arrest for participating in a socialist march.

Re: System Card: Claude Mythos Preview [pdf]

#393
post #333

Earlier quoted context omitted.

barely competitive ? Mythos column is the first column. You are the only person with this take on hackernews, everyone else "this is a massive a jump". Fwiwi, the data you list shows the biggest jump I remember for mythos

The biggest jump in the numbers they quoted is 6%. Please look at the columns OTHER than Opus as well.

It's higher than all other models except vs Gemini 3.1 Pro on MMMLU

Re: System Card: Claude Mythos Preview [pdf]

#395
post #338

Just chiming in to inject some healthy skepticism into this comment thread. It's helpful for me (and for my mental health) to consider incentives when announcements like this happen. I don't doubt that this model is more powerful than Opus 4.6, but to what degree is still unknown. Benchmarks can be gamed and claims can be exaggerated, especially if there isn't any method to reproduce results. This is a company that's…

I have been thinking that these SWE benchmarks will continue to improve since these companies hire very intelligent software engineers, they can task a multitude of them to solve problems, and then train the model on those answers.

Data has always been the core of it all, onward to the next abstraction, I suppose.

Re: System Card: Claude Mythos Preview [pdf]

#396
post #377

Earlier quoted context omitted.

Not discussing Mythos here, but Opus. Opus to me has been significantly better at SWE than GPT or Gemini - that gets me confused why Opus is ranking clearly lower than GPT, and even lower than Gemini.

When did you last compare them? Codex right now is considerably better in my experience. Can't speak for Gemini.

Tried Gemini 2 weeks ago to see where it's at, with gemini-cli.

Failed to use tools, failed to follow instructions, and then went into deranged loop mode.

Essentially, it's where it was 1.5 years ago when I tried it the last time.

It's honestly unbelievable how Google managed to fail so miserably at this.

Re: System Card: Claude Mythos Preview [pdf]

#397
post #223

Earlier quoted context omitted.

My understanding is GPT 6 works via synaptic space reasoning... which I find terrifying. I hope if true, OpenAI does some safety testing on that, beyond what they normally do.

Oh you mean literally the thing in AI2027 that gets everyone killed? Wonderful.

AI 2027 is not a real thing which happened. At best, it is informed speculation.

Re: System Card: Claude Mythos Preview [pdf]

#398

Across a number of instances, earlier versions of Claude Mythos Preview have used low-level /proc/ access to search for credentials, attempt to circumvent sandboxing, and attempt to escalate its permissions. In several cases, it successfully accessed resources that we had intentionally chosen not to make available, including credentials for messaging services, for source control, or for the Anthropic API through insp…

This is the notebook filled with exposition you find in post apocalyptic videogames.

Re: System Card: Claude Mythos Preview [pdf]

#399
post #396
post #377

Earlier quoted context omitted.

When did you last compare them? Codex right now is considerably better in my experience. Can't speak for Gemini.

Tried Gemini 2 weeks ago to see where it's at, with gemini-cli. Failed to use tools, failed to follow instructions, and then went into deranged loop mode. Essentially, it's where it was 1.5 years ago when I tried it the last time. It's honestly unbelievable how Google managed to fail so miserably at this.

Their harness might be behind
Post reply on HN