Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

251–260 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#251
post #180

Earlier quoted context omitted.

Describing providing a highly valuable service for money as `rent seeking` is pretty wild.

It could be, formally, if they have a monopoly. However, I’m tempted to compare to GitHub: if I join a new company, I will ask to be included to their GitHub account without hesitation. I couldn’t possibly imagine they wouldn’t have one. What makes the cost of that subscription reasonable is not just GitHub’s fear a crowd with pitchforks showing to their office, by also the fact that a possible answer to my non-quest…

> It could be, formally, if they have a monopoly.

you have 2 labs at the forefront (Anthropic/OpenAI), Google closely behind, xAI/Meta/half a dozen chinese companies all within 6-12 months. There is plenty of competition and price of equally intelligent tokens rapidly drop whenever a new intelligence level is achieved.

Unless the leading company uses a model to nefariously take over or neutralize another company, I don't really see a monopoly happening in the next 3 years.

Re: System Card: Claude Mythos Preview [pdf]

#252
post #53

"Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available." Disappointing that AGI will be for the powerful only. We are heading for an AI dystopia of Sci-Fi novels.

Expected outcome. Nick Land and the CCRU have explored how capitalism operationalizes science fiction (distilled in the concept of Hyperstition). Viewed through this lens, prices encode "distributed SF narratives." [0]

[0] Nick Land (1995). No Future in Fanged Noumena: Collected Writings 1987-2007, Urbanomic, p. 396.

Re: System Card: Claude Mythos Preview [pdf]

#253
post #228

Earlier quoted context omitted.

Translation: yay, more paternalism.

Anthropic always goes on and on about how their models are world changing and super dangerous like every single time they make something new they say its going to rewrite everything and scary lmao funny because they do it every time like clockwork acting like their ai is a thunderstorm coming to wipe out the world

If there are advancements, they have to be described somehow.

What if the capability advancements are real and they warrant a higher level of concern or attention?

Are we just going to automatically dismiss them because "bro, you're blowing it up too much"

Either way these improvements to capabilities are ratcheting along at about the pace that many people were expecting (and were right to expect). There is no apparent reason they will stop ratcheting along any time soon.

The rational approach is probably to start behaving as if models that are as capable as Anthropic says this one is do actually exist (even if you don't believe them on this one). The capabilities will eventually arrive, most likely sooner than we all think, and you don't want to be caught with your pants down.

Re: System Card: Claude Mythos Preview [pdf]

#254
post #230

Earlier quoted context omitted.

Yeah, need some good RE benchmarks for the LLMs. :) RE is very interesting problem. A lot more that SWE can be RE'd. I've found the LLMs are reluctant to assist, though you can workaround.

What is RE in this context?

Reverse engineering

Re: System Card: Claude Mythos Preview [pdf]

#255

I've long maintained that the real indicator that AGI is imminent is that public availability stops being a thing. If you truly believed you had a superhuman, godlike mind in your thrall, renting it out for $20/month would be the last thing you would choose to do with it.

You have to recoup your training costs though? But I’m sure you would have better option than renting it to the general public if you indeed have a perfected AI

Re: System Card: Claude Mythos Preview [pdf]

#256

Interesting reading. They are still focusing on "catastrophic risks" related to chemical and biological weapons production; or misaligned models wreaking havoc. But they are not addressing the elephant in the room: * Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement.

The unemployment rate in the US is whatever the Fed wants it to be, and isn't a function of available technology.

Re: System Card: Claude Mythos Preview [pdf]

#257
post #18

At what point do these companies stop releasing models and just use them to bootstrap AGI for themselves?

I think it is naive to think the government (US or China most probably) will just let some random company control something so powerful and dangerous.

Re: System Card: Claude Mythos Preview [pdf]

#258
post #228

Earlier quoted context omitted.

Anthropic always goes on and on about how their models are world changing and super dangerous like every single time they make something new they say its going to rewrite everything and scary lmao funny because they do it every time like clockwork acting like their ai is a thunderstorm coming to wipe out the world

If there are advancements, they have to be described somehow. What if the capability advancements are real and they warrant a higher level of concern or attention? Are we just going to automatically dismiss them because "bro, you're blowing it up too much" Either way these improvements to capabilities are ratcheting along at about the pace that many people were expecting (and were right to expect). There is no appare…

I believe advancements sure. But it is a very boy who cried wolf situation for some of these. There are other companies that behave less in this way, Antrhopic seem very unique in that they love making every single release a world ender

Re: System Card: Claude Mythos Preview [pdf]

#259
post #121
post #18

At what point do these companies stop releasing models and just use them to bootstrap AGI for themselves?

Can LLMs be AGI at all?

My understanding is no. But the definition of AGI isn’t that well defined and has been evolving, making the assessment pretty much impossible

Re: System Card: Claude Mythos Preview [pdf]

#260

Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…

> Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) > Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% > GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% > MMMLU: 92.7% / 91.1% / — / 92.6–93.6% > USAMO: 97.6% / 42.3% / 95.2% / 74.4% > OSWorld: 79.6% / 72.7% / 75.0% / — Given that for a number of these benchmarks, it seems to be barely competitive with the previous gen Opus 4.6 or GPT-5.4, I do…

> Given that for a number of these benchmarks, it seems to be barely competitive with the previous gen

We're not reading the same numbers I think. Compared to Opus 4.6, it's a big jump nearly in every single bench GP posted. They're "only" catching up to Google's Gemini on GPQA and MMMLU but they're still beating their own Opus 4.6 results on these two.

This sounds like a much better model than Opus 4.6.

Post reply on HN