Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

281–290 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#281

I've long maintained that the real indicator that AGI is imminent is that public availability stops being a thing. If you truly believed you had a superhuman, godlike mind in your thrall, renting it out for $20/month would be the last thing you would choose to do with it.

Simpler explanation : they don't have enough GPUs to release this much larger model.

Re: System Card: Claude Mythos Preview [pdf]

#284
post #63

> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. Shame. Back to business as usual then.

I for one applaud them for being cautious.

Cautious for what? Unchecked doomerism? Just release the damn models. Do it in phases, roll it out slowly if they are so damn worried about "safety".

The real reason they aren't releasing it yet is probably it eats TPU for breakfast, lunch, and dinner and inbetween.

Re: System Card: Claude Mythos Preview [pdf]

#285

Combined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2%…

damn... ok that's impressive.

Re: System Card: Claude Mythos Preview [pdf]

#286
So what changed? They are surely not getting new data to train with, what is the change in architecture that caused this? Do we not know anything about this model? My fear is Anthropic cannot be the only one that achieved it, OpenAI, Gemini and even the Chinese companies see this and probably achieved it too. At which point not releasing will become moot.

Re: System Card: Claude Mythos Preview [pdf]

#287

Earlier quoted context omitted.

> Given that for a number of these benchmarks, it seems to be barely competitive with the previous gen We're not reading the same numbers I think. Compared to Opus 4.6, it's a big jump nearly in every single bench GP posted. They're "only" catching up to Google's Gemini on GPQA and MMMLU but they're still beating their own Opus 4.6 results on these two. This sounds like a much better model than Opus 4.6.

> We're not reading the same numbers I think. We must not be. That's why I listed out the ones where it is barely competitive from @babelfish's table, which itself is extracted from Pg 186 & 187 of the System Card, which has the comparison with Opus 4.6, GPT 5.4 and Gemini 3.1 Pro. Sure, it may be better than Opus 4.6 on some of those, but barely achieves a small increase over GPT-5.4 on the ones I called out.

barely competitive ? Mythos column is the first column.

You are the only person with this take on hackernews, everyone else "this is a massive a jump". Fwiwi, the data you list shows the biggest jump I remember for mythos

Re: System Card: Claude Mythos Preview [pdf]

#288
post #223
post #26

Earlier quoted context omitted.

There's speculation that next Tuesday will be a big day for OpenAI and possibly GPT 6. Anthropic showed their hand today.

My understanding is GPT 6 works via synaptic space reasoning... which I find terrifying. I hope if true, OpenAI does some safety testing on that, beyond what they normally do.

Likely an improvement on:

> We study a novel language model architecture that is capable of scaling test-time computation by implicitly reasoning in latent space. Our model works by iterating a recurrent block, thereby unrolling to arbitrary depth at test-time. This stands in contrast to mainstream reasoning models that scale up compute by producing more tokens. Unlike approaches based on chain-of-thought, our approach does not require any specialized training data, can work with small context windows, and can capture types of reasoning that are not easily represented in words. We scale a proof-of-concept model to 3.5 billion parameters and 800 billion tokens. We show that the resulting model can improve its performance on reasoning benchmarks, sometimes dramatically, up to a computation load equivalent to 50 billion parameters.

https://arxiv.org/abs/2502.05171>

Re: System Card: Claude Mythos Preview [pdf]

#289
post #231

Earlier quoted context omitted.

I've been increasingly "freaking out" since about 3 - 4 years ago and it seems that the pessimistic scenario is materializing. It looks like it will be over for software engineers in a not so distant future. In January 2025 I said that I expect software engineers to be replaced in 2 years (pessimistic) to 5 years (optimistic). Right now I'm guessing 1 to 3 years.

I assure you it will soon become very clear that mass job losses are one of the least concerning side effects of developing the magic "everything that can plausibly been done within the constraints of physics is now possible" machine. We're opening a can of worms which I don't think most people have the imagination to understand the horrors of.

Do you have any sources I could read to better understand your concern?

Re: System Card: Claude Mythos Preview [pdf]

#290
post #172

Earlier quoted context omitted.

They could help you build an AGI if someone else has already built AGI and published it on GitHub.

I see this statement all the time and it's just strange to me. Yes, the LLMs struggle to form unique ideas - but so do we. Most advancements in human history are incremental. Built on the shoulders of millions of other incremental advancements. What i don't understand is how we quantify our ability to actually create something novel, truly and uniquely novel. We're discussing the LLMs inability to do that, yet i don'…

I really love Andrej Karpathy's take on LLMs as being instead of intelligence or sentience, a kind of cortical tissue.

It should be clear from working with LLMs over the past 4 years that they are not consciousness.

Andrej's appearance on the Dwarkesh podcast is great.

Post reply on HN