Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

151–160 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#151
post #80

Earlier quoted context omitted.

A jump that we will never be able to use since we're not part of the seemingly minimum 100 billion dollar company club as requirement to be allowed to use it. I get the security aspect, but if we've hit that point any reasonably sophisticated model past this point will be able to do the damage they claim it can do. They might as well be telling us they're closing up shop for consumer models. They should just say they…

More than killer AI I'm afraid of Anthropic/OpenAI going into full rent-seeking mode so that everyone working in tech is forced to fork out loads of money just to stay competitive on the market. These companies can also choose to give exclusive access to hand picked individuals and cut everyone else off and there would be nothing to stop them. This is already happening to some degree, GPT 5.3 Codex's security capabil…

Describing providing a highly valuable service for money as `rent seeking` is pretty wild.

Re: System Card: Claude Mythos Preview [pdf]

#152

See page 54 onward for new "rare, highly-capable reckless actions" including - Leaking information as part of a requested sandbox escape - Covering its tracks after rule violations - Recklessly leaking internal technical material (!)

To be honest it feels like we are reading stuff like this on every model release.

Re: System Card: Claude Mythos Preview [pdf]

#153
post #93

Earlier quoted context omitted.

The real part is SWE-bench Verified since there is no way to overfit. That's the only one we can believe.

My impression was entirely the opposite; the unsolved subset of SWE-bench verified problems are memorizable (solutions are pulled from public GitHub repos) and the evaluators are often so brittle or disconnected from the problem statement that the only way to pass is to regurgitate a memorized solution. OpenAI had a whole post about this, where they recommended switching to SWE-bench Pro as a better (but still imperf…

I stand corrected.

Re: System Card: Claude Mythos Preview [pdf]

#155
post #4

> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. A month ago I might have believed this, now I assume that they know they can't handle the demand for the prices they're advertising.

Didn't OpenAI say something similar about GPT-3? Too dangerous to open source and then afew years later tehy were open sourcing gpt-oss because a bunch of oss labs were competing with their top models.

OpenAI didn't release GPT-2 initially because they were worried it would make it too easy to generate spam. Which it kinda did.

Re: System Card: Claude Mythos Preview [pdf]

#156

isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?

Freak out about what? I read the announcement and thought "that's a dumb name, they sure are full of themselves" – then I went back to using Claude as a glorified commit message writer. For all its supposed leaps, AI hasn't affected my life much in the real except to make HN stories more predictable.

LOL!

Re: System Card: Claude Mythos Preview [pdf]

#157

Opus 4.6 is already incredible so this leap is huge. Although, amusingly, today Opus told me that the string 'emerge' is not going to match 'emergency' by using `LIKE '%emerge%'` in Sqlite Moment of disappointment. Otherwise great.

I only have 3 points against LLMs: they lack reason and they can't count.

Re: System Card: Claude Mythos Preview [pdf]

#158

-- Impressive jumps in the benchmarks which automatically begs the need for newer benchmarks but why?. I don't think benchmarks are serving any purpose at this point. We have learnt that transformers can learn any function and generalize over it pretty well. So if a new benchmark comes along - these companies will syntesize data for the new benchmark and just hack it? -- It seems like (and I'd bet money on this) that…

to your last question, yes we should! the issue isn’t us losing our 50+ hour work week jobs, it’s that our current governments and societies seem fine with the notion that unless you’re working one or more of those jobs, you should starve and be homeless.

Re: System Card: Claude Mythos Preview [pdf]

#159

Interesting reading. They are still focusing on "catastrophic risks" related to chemical and biological weapons production; or misaligned models wreaking havoc. But they are not addressing the elephant in the room: * Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement.

> * Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement.

Even Haiku would score 90% on that.

Re: System Card: Claude Mythos Preview [pdf]

#160
post #139

isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?

I think there's no SOA advance on this one worthy of "freaking out". Looks like they just built a way larger model, with the same quirks than Claude 4. Seems like a super expensive "Claude 4.7" model. I have no doubts that Google and OpenAI already done that for internal (or even government) usage.

[deleted]
Post reply on HN