Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

331–340 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#331
post #208

> Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin. We believe that it does not have any significant coherent misaligned goals, and its character traits in typical conversations closely follow the goals we laid out in our constitution. Even so, we believe that it likely poses the greatest alignment-related risk of any…

Alignment “appearing” better as model capabilities increase scares the shit out of me, tbh.

Conversely: in humans, intelligence is inversely correlated with crime.

It doesn't go to zero, however!

Re: System Card: Claude Mythos Preview [pdf]

#332

So what changed? They are surely not getting new data to train with, what is the change in architecture that caused this? Do we not know anything about this model? My fear is Anthropic cannot be the only one that achieved it, OpenAI, Gemini and even the Chinese companies see this and probably achieved it too. At which point not releasing will become moot.

New pre train?

Re: System Card: Claude Mythos Preview [pdf]

#333

Earlier quoted context omitted.

> We're not reading the same numbers I think. We must not be. That's why I listed out the ones where it is barely competitive from @babelfish's table, which itself is extracted from Pg 186 & 187 of the System Card, which has the comparison with Opus 4.6, GPT 5.4 and Gemini 3.1 Pro. Sure, it may be better than Opus 4.6 on some of those, but barely achieves a small increase over GPT-5.4 on the ones I called out.

barely competitive ? Mythos column is the first column. You are the only person with this take on hackernews, everyone else "this is a massive a jump". Fwiwi, the data you list shows the biggest jump I remember for mythos

The biggest jump in the numbers they quoted is 6%.

Please look at the columns OTHER than Opus as well.

Re: System Card: Claude Mythos Preview [pdf]

#334

Earlier quoted context omitted.

Messy in a way that would affect you?

"Internet no longer viable" would affect everyone, probably

The only thing preventing this today is cost, not capability. As costs come down over the next 5 years, the idea that the internet was once dominated by people will seem quaint.

Re: System Card: Claude Mythos Preview [pdf]

#336

So what changed? They are surely not getting new data to train with, what is the change in architecture that caused this? Do we not know anything about this model? My fear is Anthropic cannot be the only one that achieved it, OpenAI, Gemini and even the Chinese companies see this and probably achieved it too. At which point not releasing will become moot.

Well the important thing is they have a lot more data of people actually using their models. They have read billions more lines of private repos and implemented millions of patches, all of which is feeding into the newer models. More importantly it understand what behaviour people tend to appreciate and what changes are more likely to get approved. This real world usage data is invaluable.

Exactly. As Claude increases in popularity, their available training data also increases. I'd guess Anthropic has the most expansive swe training data as of now, if not close. Considering how quickly Claude is penetrating, I expect their lead to grow quickly.

Re: System Card: Claude Mythos Preview [pdf]

#337
post #290

Earlier quoted context omitted.

I see this statement all the time and it's just strange to me. Yes, the LLMs struggle to form unique ideas - but so do we. Most advancements in human history are incremental. Built on the shoulders of millions of other incremental advancements. What i don't understand is how we quantify our ability to actually create something novel, truly and uniquely novel. We're discussing the LLMs inability to do that, yet i don'…

I really love Andrej Karpathy's take on LLMs as being instead of intelligence or sentience, a kind of cortical tissue. It should be clear from working with LLMs over the past 4 years that they are not consciousness. Andrej's appearance on the Dwarkesh podcast is great.

To be clear i agree with you, my question is more pointed at us - i'm not sure we have a good understanding of conciousness, nor that we are as we seem. Given how prone to hallucinations we are, how our subtle hormones can drastically alter what we perceive as our intelligence, self identity, etc.

I'm not convinced LLMs are anything amazing in their current form, but i suspect they'll push a self reflection on us.

But clearly i think humans are far more Input-Output than the average person. I'm also not educated on the subject, so what do i know hah.

Re: System Card: Claude Mythos Preview [pdf]

#338
Just chiming in to inject some healthy skepticism into this comment thread. It's helpful for me (and for my mental health) to consider incentives when announcements like this happen.

I don't doubt that this model is more powerful than Opus 4.6, but to what degree is still unknown. Benchmarks can be gamed and claims can be exaggerated, especially if there isn't any method to reproduce results.

This is a company that's battling it out with a number of other well-funded and extremely capable competitors. What they've done so far is remarkable, but at the end of the day they want to win this race. They also have an upcoming IPO.

Scare-mongering like this is Anthropic's bread and butter, they're extremely good at it. They do it in a subtle and almost tasteful way sometimes. Their position as the respectable AI outfit that caters to enterprise gives them good footing to do it, too.

Re: System Card: Claude Mythos Preview [pdf]

#339
post #333

Earlier quoted context omitted.

barely competitive ? Mythos column is the first column. You are the only person with this take on hackernews, everyone else "this is a massive a jump". Fwiwi, the data you list shows the biggest jump I remember for mythos

The biggest jump in the numbers they quoted is 6%. Please look at the columns OTHER than Opus as well.

> Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro)

> Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5%

> USAMO: 97.6% / 42.3% / 95.2% / 74.4%

> The biggest jump in the numbers they quoted is 6%.

Just in the numbers you quoted, thats a 16.6% jump in terminal-bench and a 55.3% absolute increase in USAMO over their previous Opus 4.6 model.

Re: System Card: Claude Mythos Preview [pdf]

#340
post #267

Earlier quoted context omitted.

to your last question, yes we should! the issue isn’t us losing our 50+ hour work week jobs, it’s that our current governments and societies seem fine with the notion that unless you’re working one or more of those jobs, you should starve and be homeless.

This is a theory I can't support well beyond hypothesising about what a post-employment democracy might look like, but I strongly suspect democracy doesn't work in a world where voters neither hold any significant collective might and are not producing any significant wealth. Democracies work because people collectively have power, in previous centuries that was partly collective physical might, but in recent years i…

The only way to avoid corruption is to take power out of human hands. Historically, this had meant shifting the power to markets, but when markets cease to function in a way that allows people to feed themselves, we will need to find another way.

I hate to say it, but gold bugs, crypto bros, and AI governance people might be onto something.

Post reply on HN