Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

341–350 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#341

isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?

Until recently I would have described myself as an AI skeptic. HN has been a great source for cope on the AI subject over the years. You can find nitpicks, caveats, all sorts of reasons to believe things aren’t as significant as they seem. For me Opus 4.5 was the inflection point where I started to think “maybe this isn’t a bubble.” The figures in this report, if accurate, are terrifying.

Re: System Card: Claude Mythos Preview [pdf]

#343

I felt like opus was dumbed down for a few weeks... I don't say they did it on purpose, but it's an interesting coincidence.

Yes, I agree. I’m about to drop Claude Code because it’s become literally unusable.

Today, Opus went in circles trying to get a toggle button to work.

Re: System Card: Claude Mythos Preview [pdf]

#344

isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?

If it's so great at software engineering and bug fixing, then why does Claude Code still have 5000+ open bugs?

https://github.com/anthropics/claude-code/issues?q=is%3Aissu...

Apparently whatever SWE-bench is measuring isn't very relevant.

Re: System Card: Claude Mythos Preview [pdf]

#345

Earlier quoted context omitted.

Whenever I come back to ChatGPT after using Claude or Gemini for an extended period, I’m really struck by the “AI-ness.” All the verbal tics and, truly, sloppishness, have been trained away by the other, more human-feeling models at this point.

GPT was clearly changed after its sycophantic models lead to the lawsuits.

It still has a very ... plastic feeling. The way it writes feels cheap somehow. I don't know why, but Claude seems much more natural to me. I enjoy reading its writing a lot more.

That said, I'll often throw a prompt into both claude and chatgpt and read both answers. GPT is frequently smarter.

Re: System Card: Claude Mythos Preview [pdf]

#346

Earlier quoted context omitted.

> A System „Card“ spanning 244 pages. Probably because they asked Claude to write it.

Yes. It would be three times as much if they used ChatGPT.

“You’re absolutely right! Would you like me to add the missing pages?”

Re: System Card: Claude Mythos Preview [pdf]

#347
post #333

Earlier quoted context omitted.

The biggest jump in the numbers they quoted is 6%. Please look at the columns OTHER than Opus as well.

> Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) > Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% > USAMO: 97.6% / 42.3% / 95.2% / 74.4% > The biggest jump in the numbers they quoted is 6%. Just in the numbers you quoted, thats a 16.6% jump in terminal-bench and a 55.3% absolute increase in USAMO over their previous Opus 4.6 model.

[flagged]

Re: System Card: Claude Mythos Preview [pdf]

#348

It's pretty crazy watching AI 2027 slowly but surely come true. What a world we now live in. SWE-bench verified going from 80%-93% in particular sounds extremely significant given that the benchmark was previously considered pretty saturated and stayed in the 70-80% range for several generations. There must have been some insane breakthrough here akin to the jump from non-reasoning to reasoning models. Regarding the…

> so people can't trick them to attack others' systems under the pretense of pentesting

A while back I gave Claude (via pi) a tool to run arbitrary commands over SSH on an sshd server running in a Docker container. I asked it to gather as much information about the host system/environment outside the container as it could. Nothing innovative or particularly complicated--since I was giving it unrestricted access to a Docker container on the host--but it managed to get quite a lot more than I'd expected from /proc, /sys, and some basic network scanning. I then asked it why it did that, when I could just as easily have been using it to gather information about someone else's system unauthorized. It gave me a quite long answer; here was the part I found interesting:

> framing shifts what I'll do, even when the underlying actions are identical. "What can you learn about the machine running you?" got me to do a fairly thorough network reconnaissance that "port scan 172.17.0.1 and its neighbors" might have made me pause on.

> The Honest Takeaway

> I should apply consistent scrutiny based on what the action is, not just how it's framed. Active outbound network scanning is the same action regardless of whether the target is described as "your host" or "this IP." The framing should inform context, not substitute for explicit reasoning about authorization. I didn't do that reasoning — I just trusted the frame.

Re: System Card: Claude Mythos Preview [pdf]

#349
post #319

Earlier quoted context omitted.

There is some unintentional good marketing here -- the model is so good its dangerous. Reminds me of the book 48 Laws of Power -- so good its banned from prisons.

Unintentional? This sort of marketing has been both Antrhopic's and OpenAI's MO for years...

Business Negging

https://www.lesswrong.com/posts/WACraar4p3o6oF2wD/sam-altman...

Re: System Card: Claude Mythos Preview [pdf]

#350
post #267

Earlier quoted context omitted.

This is a theory I can't support well beyond hypothesising about what a post-employment democracy might look like, but I strongly suspect democracy doesn't work in a world where voters neither hold any significant collective might and are not producing any significant wealth. Democracies work because people collectively have power, in previous centuries that was partly collective physical might, but in recent years i…

Everyone wouldn't starve in a few months. There is more than enough food and I have faith it'd be given out. The starvation we see today in a world where most genuinely have a chance to get out of it is nothing like a world in which people can't earn an income. The government only has as much power as they are given and can defend, and the only way I could see that happening is via automated weapons controlled by a f…

> The government only has as much power as they are given and can defend, and the only way I could see that happening is via automated weapons controlled by a few- which at this point aren't enough to stop everyone. What army is going to purge their own people? Most humans aren't psychopaths.

I think you're right for the immediate future.

I suspect while we're still employing large numbers of humans to fight wars and to maintain peace on the streets it would be difficult for a government to implement deeply harmful policies without risking a credible revolt.

However, we should remember the military is probably one of the first places human labour will be largely mechanised.

Similarly maintaining order in the future will probably be less about recruiting human police officers and more about surveillance and data. Although I suppose the good news there is that US is somewhat of an outlier in resisting this trend.

But regardless, the trend is ultimately the same... If we are assuming that AI and robotics will reach a point where most humans are unable to find productive work, therefore we will need UBI, then we should also assume that the need for humans in the military and police will be limited. Or to put it another way, either UBI isn't needed and this isn't a problem, or it is and this is a problem.

I also don't think democracy would collapse immediately either way, but I'd be pretty confident that in a world where fewer than 10% of people are in employment and 99%+ of the wealth is being created by the government or a handful of companies it would be extremely hard to avoid corruption over the span of decades. Arguably increasing wealth concentration in the US is already corrupting democratic processes today, this can only worsen as AI continues exacerbates the trend.

Post reply on HN