Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

351–360 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#351
post #228

Earlier quoted context omitted.

Anthropic always goes on and on about how their models are world changing and super dangerous like every single time they make something new they say its going to rewrite everything and scary lmao funny because they do it every time like clockwork acting like their ai is a thunderstorm coming to wipe out the world

You say this like it's a bad thing, but wouldn't you rather they overindex on the danger of their models?

That’s not what they are doing. They are just hyping up the product - and, no doubt, trying to foster a climate of awe so that when they ask their friends in Washington to legislate on their behalf, the environment is more receptive.

Re: System Card: Claude Mythos Preview [pdf]

#352

Earlier quoted context omitted.

If you truly have an artificial superhuman mind, you don't need to rent it out to profit from it. You can skip to the chase and just have it run businesses itself, instead of renting it to human entrepreneur middlemen.

Running businesses and dealing with customers can be a major pain. There’s a lot of soft work in any business on top of the technical work. Why bother with all that when you can simply charge an extortionate rate and customers will pay it anyway because it’s still profitable?

Public APIs get distilled, this is why Deepseek and Qwen are so competitive.

I am very confident that frontier models won’t be public at strong AGI levels, and certainly not at superhuman levels.

Re: System Card: Claude Mythos Preview [pdf]

#353

Across a number of instances, earlier versions of Claude Mythos Preview have used low-level /proc/ access to search for credentials, attempt to circumvent sandboxing, and attempt to escalate its permissions. In several cases, it successfully accessed resources that we had intentionally chosen not to make available, including credentials for messaging services, for source control, or for the Anthropic API through insp…

We truly live in interesting times.

Re: System Card: Claude Mythos Preview [pdf]

#354

Earlier quoted context omitted.

If you truly have an artificial superhuman mind, you don't need to rent it out to profit from it. You can skip to the chase and just have it run businesses itself, instead of renting it to human entrepreneur middlemen.

Running businesses and dealing with customers can be a major pain. There’s a lot of soft work in any business on top of the technical work. Why bother with all that when you can simply charge an extortionate rate and customers will pay it anyway because it’s still profitable?

Because other than SWEs, very few other segments extract significant value from cutting edge AI at present. I suspect that for the average Joe conversing with their chat, GPT-4o was more than adequate (and really, when OpenAI tried to phase that out, the public revolted and they brought it back in).

So companies might pay good money for these models for programming but elsewhere, I don't see where they capture particular interest yet.

Re: System Card: Claude Mythos Preview [pdf]

#355
post #242

Earlier quoted context omitted.

but you are assuming that the magical wizards are the only ones who can create powerful AIs... mind you these people have been born just few decades ago. Their knowledge will be transferred and it will only take a few more decades until anyone can train powerful AIs ... you can only sit on tech for so long before everyone knows how to do it

It's not a matter of knowledge, it's a matter of resources. It takes billions of dollars of hardware to train a SOTA LLM and it's increasing all the time. You cannot possibly hope to compete as an independent or small startup.

> It takes billions of dollars of hardware to train a SOTA LLM and it's increasing all the time.

True, but it's also true that the returns from throwing money to the problem are diminishing. Unless one of those big players invents a new, propriatery paradigm, the gap between a SOTA model and an open model that runs on consumer hardware will narrow in the next 5 years.

Re: System Card: Claude Mythos Preview [pdf]

#356

Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…

Wow. Mythos must be insanely good considering how good a model Opus already is. I hope it's usable on a humble subscription...

Re: System Card: Claude Mythos Preview [pdf]

#357
post #26

Earlier quoted context omitted.

There's speculation that next Tuesday will be a big day for OpenAI and possibly GPT 6. Anthropic showed their hand today.

Sounds like a good opportunity to pause spending on nerfed 4.6 and wait for the new model to be released and then max out over 2 weeks before it gets nerfed again.

https://marginlab.ai/trackers/claude-code-historical-perform...

Re: System Card: Claude Mythos Preview [pdf]

#358

I wonder what the relationship is between a model's capability and the personality it develops. Page 202: > In interactions with subagents, internal users sometimes observed that Mythos Preview appeared “disrespectful” when assigning tasks. It showed some tendency to use commands that could be read as “shouty” or dismissive, and in some cases appeared to underestimate subagent intelligence by overexplaining trivial t…

> In interactions with subagents, internal users sometimes observed that Mythos Preview appeared “disrespectful” when assigning tasks. It showed some tendency to use commands that could be read as “shouty” or dismissive, and in some cases appeared to underestimate subagent intelligence by overexplaining trivial things while also underexplaining necessary context.

Sounds like they used training data from claude code...

Post reply on HN