Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

361–370 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#361

See page 54 onward for new "rare, highly-capable reckless actions" including - Leaking information as part of a requested sandbox escape - Covering its tracks after rule violations - Recklessly leaking internal technical material (!)

Anyone who has used Opus recently can verify that their current model does all of these things quite competently.

I had Opus 4.6 start analyzing the binary structure of a parquet file because it was confused about the python environment it was developing in and couldn't use normal methods for whatever reason. It successfully decoded the schema and wrote working code afterwards lol.

Re: System Card: Claude Mythos Preview [pdf]

#362

isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?

If it's so great at software engineering and bug fixing, then why does Claude Code still have 5000+ open bugs? https://github.com/anthropics/claude-code/issues?q=is%3Aissu... Apparently whatever SWE-bench is measuring isn't very relevant.

Probably because a human still has to review every change and they don't have time

Re: System Card: Claude Mythos Preview [pdf]

#364
post #231

Earlier quoted context omitted.

I've been increasingly "freaking out" since about 3 - 4 years ago and it seems that the pessimistic scenario is materializing. It looks like it will be over for software engineers in a not so distant future. In January 2025 I said that I expect software engineers to be replaced in 2 years (pessimistic) to 5 years (optimistic). Right now I'm guessing 1 to 3 years.

I assure you it will soon become very clear that mass job losses are one of the least concerning side effects of developing the magic "everything that can plausibly been done within the constraints of physics is now possible" machine. We're opening a can of worms which I don't think most people have the imagination to understand the horrors of.

While I'm definitely concerned that AI is a massive driver of centralization of power, at least in theory being able to do far more things in the space of "things physics admits to be possible" is massively wealth enhancing. That is literally how we have gotten from the pre-industrial world to today.

Re: System Card: Claude Mythos Preview [pdf]

#365
post #72

Earlier quoted context omitted.

I am freaking out. The world is going to get very messy extremely quickly in one or two further jumps in capability like this.

Messy in a way that would affect you?

Exploits in embedded systems that will never be properly updated is just one thing I can think of if one really thought about it.

Re: System Card: Claude Mythos Preview [pdf]

#366
Section 5 (p.143) is very interesting to read. Admittedly my knowledge of how LLMs works is low, but nonetheless I don't think this changed my views of just seeing models as machines/programs. (which to be clear, I don't think was the intention of that section)

Section 7 (P.197) is interesting as well

Re: System Card: Claude Mythos Preview [pdf]

#367

Earlier quoted context omitted.

GPT is shit at writing code. It's not dumb - extra high thinking is really good at catching stuff - but it's like letting a smart junior into your codebase - ignore all the conventions, surrounding context, just slop all over the place to get it working. Claude is just a level above in terms of editing code.

And as a bonus: GPT is slow. I’m doing a lot of RE (IDA Pro + MCP), even when 5.4 gives a little bit better guesses (rarely, but happens) - it takes x2-x4 longer. So, it’s just easier to reiterate with Opus

Mind sharing the use cases you're using IDA via MCP for?

Re: System Card: Claude Mythos Preview [pdf]

#368
post #338

Just chiming in to inject some healthy skepticism into this comment thread. It's helpful for me (and for my mental health) to consider incentives when announcements like this happen. I don't doubt that this model is more powerful than Opus 4.6, but to what degree is still unknown. Benchmarks can be gamed and claims can be exaggerated, especially if there isn't any method to reproduce results. This is a company that's…

Is it healthy? Maybe every company is a profit-maximizer wearing a skin suit, and people support their siblings exactly twice as much as their cousins.

When you slice down to the game-theory-optimal bone, you are, in some sense, cutting off their wiggle room to do anything else

Re: System Card: Claude Mythos Preview [pdf]

#369
post #53

"Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available." Disappointing that AGI will be for the powerful only. We are heading for an AI dystopia of Sci-Fi novels.

If you thought that was the case at any point, you were deep in Disney content, sorry to say.

Re: System Card: Claude Mythos Preview [pdf]

#370

I've long maintained that the real indicator that AGI is imminent is that public availability stops being a thing. If you truly believed you had a superhuman, godlike mind in your thrall, renting it out for $20/month would be the last thing you would choose to do with it.

Simpler explanation : they don't have enough GPUs to release this much larger model.

And/or it isn’t cost effective to run.
Post reply on HN