isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?
System Card: Claude Mythos Preview [pdf]
341–350 of 687 posts
Re: System Card: Claude Mythos Preview [pdf]
#342Re: System Card: Claude Mythos Preview [pdf]
#343I felt like opus was dumbed down for a few weeks... I don't say they did it on purpose, but it's an interesting coincidence.
Today, Opus went in circles trying to get a toggle button to work.
Re: System Card: Claude Mythos Preview [pdf]
#344isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?
https://github.com/anthropics/claude-code/issues?q=is%3Aissu...
Apparently whatever SWE-bench is measuring isn't very relevant.
Re: System Card: Claude Mythos Preview [pdf]
#345Earlier quoted context omitted.
Whenever I come back to ChatGPT after using Claude or Gemini for an extended period, I’m really struck by the “AI-ness.” All the verbal tics and, truly, sloppishness, have been trained away by the other, more human-feeling models at this point.
GPT was clearly changed after its sycophantic models lead to the lawsuits.
That said, I'll often throw a prompt into both claude and chatgpt and read both answers. GPT is frequently smarter.
Re: System Card: Claude Mythos Preview [pdf]
#346Re: System Card: Claude Mythos Preview [pdf]
#347Earlier quoted context omitted.
The biggest jump in the numbers they quoted is 6%. Please look at the columns OTHER than Opus as well.
> Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) > Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% > USAMO: 97.6% / 42.3% / 95.2% / 74.4% > The biggest jump in the numbers they quoted is 6%. Just in the numbers you quoted, thats a 16.6% jump in terminal-bench and a 55.3% absolute increase in USAMO over their previous Opus 4.6 model.
Re: System Card: Claude Mythos Preview [pdf]
#348It's pretty crazy watching AI 2027 slowly but surely come true. What a world we now live in. SWE-bench verified going from 80%-93% in particular sounds extremely significant given that the benchmark was previously considered pretty saturated and stayed in the 70-80% range for several generations. There must have been some insane breakthrough here akin to the jump from non-reasoning to reasoning models. Regarding the…
A while back I gave Claude (via pi) a tool to run arbitrary commands over SSH on an sshd server running in a Docker container. I asked it to gather as much information about the host system/environment outside the container as it could. Nothing innovative or particularly complicated--since I was giving it unrestricted access to a Docker container on the host--but it managed to get quite a lot more than I'd expected from /proc, /sys, and some basic network scanning. I then asked it why it did that, when I could just as easily have been using it to gather information about someone else's system unauthorized. It gave me a quite long answer; here was the part I found interesting:
> framing shifts what I'll do, even when the underlying actions are identical. "What can you learn about the machine running you?" got me to do a fairly thorough network reconnaissance that "port scan 172.17.0.1 and its neighbors" might have made me pause on.
> The Honest Takeaway
> I should apply consistent scrutiny based on what the action is, not just how it's framed. Active outbound network scanning is the same action regardless of whether the target is described as "your host" or "this IP." The framing should inform context, not substitute for explicit reasoning about authorization. I didn't do that reasoning — I just trusted the frame.
Re: System Card: Claude Mythos Preview [pdf]
#349Earlier quoted context omitted.
There is some unintentional good marketing here -- the model is so good its dangerous. Reminds me of the book 48 Laws of Power -- so good its banned from prisons.
Unintentional? This sort of marketing has been both Antrhopic's and OpenAI's MO for years...
https://www.lesswrong.com/posts/WACraar4p3o6oF2wD/sam-altman...
Re: System Card: Claude Mythos Preview [pdf]
#350Earlier quoted context omitted.
This is a theory I can't support well beyond hypothesising about what a post-employment democracy might look like, but I strongly suspect democracy doesn't work in a world where voters neither hold any significant collective might and are not producing any significant wealth. Democracies work because people collectively have power, in previous centuries that was partly collective physical might, but in recent years i…
Everyone wouldn't starve in a few months. There is more than enough food and I have faith it'd be given out. The starvation we see today in a world where most genuinely have a chance to get out of it is nothing like a world in which people can't earn an income. The government only has as much power as they are given and can defend, and the only way I could see that happening is via automated weapons controlled by a f…
I think you're right for the immediate future.
I suspect while we're still employing large numbers of humans to fight wars and to maintain peace on the streets it would be difficult for a government to implement deeply harmful policies without risking a credible revolt.
However, we should remember the military is probably one of the first places human labour will be largely mechanised.
Similarly maintaining order in the future will probably be less about recruiting human police officers and more about surveillance and data. Although I suppose the good news there is that US is somewhat of an outlier in resisting this trend.
But regardless, the trend is ultimately the same... If we are assuming that AI and robotics will reach a point where most humans are unable to find productive work, therefore we will need UBI, then we should also assume that the need for humans in the military and police will be limited. Or to put it another way, either UBI isn't needed and this isn't a problem, or it is and this is a problem.
I also don't think democracy would collapse immediately either way, but I'd be pretty confident that in a world where fewer than 10% of people are in employment and 99%+ of the wealth is being created by the government or a handful of companies it would be extremely hard to avoid corruption over the span of decades. Arguably increasing wealth concentration in the US is already corrupting democratic processes today, this can only worsen as AI continues exacerbates the trend.