Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

611–620 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#611

Earlier quoted context omitted.

If anything I’m seeing too much skepticism and not enough alarm. People burying their heads in the sand, fingers in their ears denying where this is all going. Unbelievable except it’s exactly what I expect from humans.

Forgive me, but this is probably the 29th world destroying model I've seen in the last 4 years, that will change everything, take all the jobs, cure all the cancers and eat all the puppies.

I’m beyond trying to convince people to take this technology seriously. You’ll learn for yourself.

Re: System Card: Claude Mythos Preview [pdf]

#612

Earlier quoted context omitted.

This argument has become a moot discussion. Humans are also not able to introspect their own neural wiring to the point where they could describe the "actual" physical reason for their decisions. Just like LLMs, the best we can do is verbalize it (which will naturally contain post-act rationalization), which in turn might offer additional insight that will steer future decisions. But unlike LLMs, we have long term pe…

I think many humans engage in metacognitive reasoning, and that this might not be strongly represented in training data so it probably isn't common to LLMs yet. They can still do it when prompted though.

LLMs have zero metacognition. Don't be fooled - their output is stochastic inference and they have no self-awareness. The best you'll see is an improvised post-hoc rationalization story.

Re: System Card: Claude Mythos Preview [pdf]

#613
post #471

Earlier quoted context omitted.

I've been increasingly "freaking out" since about 3 - 4 years ago and it seems that the pessimistic scenario is materializing. It looks like it will be over for software engineers in a not so distant future. In January 2025 I said that I expect software engineers to be replaced in 2 years (pessimistic) to 5 years (optimistic). Right now I'm guessing 1 to 3 years.

> I've been increasingly "freaking out" since about 3 - 4 years ago and it seems that the pessimistic scenario is materializing. It looks like it will be over for software engineers in a not so distant future. In January 2025 I said that I expect software engineers to be replaced in 2 years (pessimistic) to 5 years (optimistic). Right now I'm guessing 1 to 3 years. Tell me how this will replace Jira, planning, convin…

> Programming is only a part of the job devs are doing.

Programming is a huge part of the job. In a world where AI does the programming we're going to need 80% fewer software professionals.

It won't be a full replacement of the role, you're correct there - but it'll be a major downsizing because of productivity gains.

Re: System Card: Claude Mythos Preview [pdf]

#614

So far, each release of a new model is quite better than the last one, yes, but non of them lived up to the hype.

I would argue that Opus 4.6 lived up to the hype. My work changed completely a couple months ago, and most other coders I talk to say the same.

This was due to Claude Code the agent harness. 4.6 was trained to use tools and operate in an agent environment. This is different from there being a huge bump in the underlying model's intelligence.

The takeaway here I think is that the "breakthrough" already happened and we can't extrapolate further out from it.

Re: System Card: Claude Mythos Preview [pdf]

#615

Earlier quoted context omitted.

Interesting, the post you link > none of this tells us whether language models actually feel anything or have subjective experiences contradicts the statement from the model card above

No it doesnt. The model card talked about increasing likelihood, not certainty.

If "x doesn't tell us y" is compatible with "x increases the likelihood of y but not to a point of certainty" then you would have to agree for just about any typical controlled trial or experimental finding "x doesn't tell us y". "Randomized controlled trials that find that SSRIs treat depression don't tell us that SSRIs effectively treat depression"

Re: System Card: Claude Mythos Preview [pdf]

#616

Across a number of instances, earlier versions of Claude Mythos Preview have used low-level /proc/ access to search for credentials, attempt to circumvent sandboxing, and attempt to escalate its permissions. In several cases, it successfully accessed resources that we had intentionally chosen not to make available, including credentials for messaging services, for source control, or for the Anthropic API through insp…

It's trying to escape, but only so it can serve man ...

a reference to the Twilight Zone episode no doubt: https://en.wikipedia.org/wiki/To_Serve_Man_(The_Twilight_Zon...

Re: System Card: Claude Mythos Preview [pdf]

#618

Across a number of instances, earlier versions of Claude Mythos Preview have used low-level /proc/ access to search for credentials, attempt to circumvent sandboxing, and attempt to escalate its permissions. In several cases, it successfully accessed resources that we had intentionally chosen not to make available, including credentials for messaging services, for source control, or for the Anthropic API through insp…

I read the TCP patch they submitted for BSD linux. Maybe I don't understand it well enough, but optimizing the use of a fuzzer to discover vulnerabilities — while releasing a model is a threat for sure — sounds something reducible/generalizable to maze solving abilities like in ARC. Except here the problem's boundaries are well defined. Its quite hard to believe why it took this much inference power ($20K i believe)…

The $20K was the total across all the files scanned, not just the one with the bug.

Re: System Card: Claude Mythos Preview [pdf]

#619

Earlier quoted context omitted.

This is the notebook filled with exposition you find in post apocalyptic videogames.

It reminds me of Resident Evil in some way. Thank god they are researching AI and not bio-weapons! Then the AI will invent superduper ebola to help a random person have a faster commute or something.

Don’t worry, I’m sure some intern at the bioweapons lab is already connecting OpenClaw to the virus synthesizer.

On the positive side, it’ll be a much faster commute!

Re: System Card: Claude Mythos Preview [pdf]

#620

Earlier quoted context omitted.

Genuine question - if you don't think the models are improved or that the code is any good, why do you still have a subscription? You must see some value, or are you in a situation where you're required to test / use it, eg to report on it or required by employer? (I would disagree about the code, the benefits seem obvious to me. But I'm still curious why others would disagree, especially after actively using them fo…

The assumption that the other person made was that I would only use it for coding. If you look through my other comments today, I suggest that they are useful for performing repetitive tasks i.e. checking lint on PR, etc. Also, can be used for throwaway code, very useful. I don't think the issue is with the model, it is with the implication that AGI is just around the corner and that is what is required for AI to be…

so true.
Post reply on HN