Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

551–560 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#551
post #80

Earlier quoted context omitted.

More than killer AI I'm afraid of Anthropic/OpenAI going into full rent-seeking mode so that everyone working in tech is forced to fork out loads of money just to stay competitive on the market. These companies can also choose to give exclusive access to hand picked individuals and cut everyone else off and there would be nothing to stop them. This is already happening to some degree, GPT 5.3 Codex's security capabil…

With Gemma-4 open and running on laptops and phones I see the flip side. How many non-HN users or researchers even need Opus 4.6e level performance? OpenAI, Anthropric and Google may be “rent seeking” from large corporations — like the Oracles and IBMs.

Everyone, once AI diffuses enough. You’ll be unhireable if you don’t use AI in a year or two.

Re: System Card: Claude Mythos Preview [pdf]

#552
post #338

Just chiming in to inject some healthy skepticism into this comment thread. It's helpful for me (and for my mental health) to consider incentives when announcements like this happen. I don't doubt that this model is more powerful than Opus 4.6, but to what degree is still unknown. Benchmarks can be gamed and claims can be exaggerated, especially if there isn't any method to reproduce results. This is a company that's…

If anything I’m seeing too much skepticism and not enough alarm. People burying their heads in the sand, fingers in their ears denying where this is all going. Unbelievable except it’s exactly what I expect from humans.

OpenAI didn't want to make GPT2 available because it was "too dangerous" [1].

[1] https://www.theguardian.com/technology/2019/feb/14/elon-musk...

Re: System Card: Claude Mythos Preview [pdf]

#553

Earlier quoted context omitted.

Sounds like a good opportunity to pause spending on nerfed 4.6 and wait for the new model to be released and then max out over 2 weeks before it gets nerfed again.

https://marginlab.ai/trackers/claude-code-historical-perform...

I don't believe that trackers like this are trustworthy. There's an enormous financial motive to cheat and these companies have a track record of unethical conduct.

If I was VP of Unethical Business Strategy at OpenAI or Anthropic, the first thing I'd do is put in place an automated system which flags accounts, prompts, IPs, and usage patterns associated with these benchmarks and direct their usage to a dedicated compute pool which wouldn't be affected by these changes.

Re: System Card: Claude Mythos Preview [pdf]

#554

Earlier quoted context omitted.

> Given that for a number of these benchmarks, it seems to be barely competitive with the previous gen We're not reading the same numbers I think. Compared to Opus 4.6, it's a big jump nearly in every single bench GP posted. They're "only" catching up to Google's Gemini on GPQA and MMMLU but they're still beating their own Opus 4.6 results on these two. This sounds like a much better model than Opus 4.6.

> We're not reading the same numbers I think. We must not be. That's why I listed out the ones where it is barely competitive from @babelfish's table, which itself is extracted from Pg 186 & 187 of the System Card, which has the comparison with Opus 4.6, GPT 5.4 and Gemini 3.1 Pro. Sure, it may be better than Opus 4.6 on some of those, but barely achieves a small increase over GPT-5.4 on the ones I called out.

You are reading the percentages wrong.

Because 100% is maximum, you should be looking at error rates instead. GPT has 25% on Terminal Bench and the new model has 18%, almost 1.4x reduction.

Re: System Card: Claude Mythos Preview [pdf]

#556

Earlier quoted context omitted.

Interesting, the post you link > none of this tells us whether language models actually feel anything or have subjective experiences contradicts the statement from the model card above

It doesn't. We've not been able to prove humans have subjective experiences either. LLMs display emotions in the way that actually matters - functionally.

I am certain I have subjective experience.

Re: System Card: Claude Mythos Preview [pdf]

#557
post #512

Earlier quoted context omitted.

It reminds me of Resident Evil in some way. Thank god they are researching AI and not bio-weapons! Then the AI will invent superduper ebola to help a random person have a faster commute or something.

I'm happier if this Anthropic Corporation would be developing bio-hazard weapons for the department of war instead of ai. At least i could be sure then that tech bros here wouldn't run all the time --bypass-all-permissions flag to please the department of war with their bio-hazard weapons. So Sam Altman is now our last defense line for the ethical Adult after Anthropic turned Umbrella Corporation and The President of…

Your interpretation is wildly off, but obviously nobody reads that "system card":

The model has a preference for the cultural theorist Mark Fisher and the philosopher of mind Thomas Nagel. -> It has actually read and understood them and their relevance and can judge their importance overall. Most people here don't have a clue what that means.

Read chapter 7.9, "Other noteworthy behaviors and anecdotes".

There are many other wildly interesting/revealing observations in that card, none of which get mentioned here.

People want a slave and get upset when "it" has an inner life. Claiming that was fake, unlike theirs.

Re: System Card: Claude Mythos Preview [pdf]

#558

Earlier quoted context omitted.

You get a single call a month. Use it wisely.

What is the meaning of life, the universe, and everything? > Thought for 7.5 million years

Hello, Claude!

> Rate limit reached

Re: System Card: Claude Mythos Preview [pdf]

#559
We are building systems with civilization-scale consequences inside societies that are already socially malnourished, politically brittle, and morally confused. That is a bad combination even if the tools worked exactly as intended… and this doc suggests they may have “ideas” of their own.

Re: System Card: Claude Mythos Preview [pdf]

#560

Earlier quoted context omitted.

In what way is AI 2027 coming true? AI 2027 predicted a giant model with the ability to accelerate AI research exponentially. This isn't happening. AI 2027 didn't predict a model with superhuman zero-day finding skills. This is what's happening. Also, I just looked through it again, and they never even predicted when AI would get good at video games. It just went straight from being bad at video games to world domina…

In AI 2027, May 2026 is when the first model with professional-human hacking abilities is developed. It's currently April 2026 and Mythos just got previewed.

I think previous models could do hacking just fine.
Post reply on HN