Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

461–470 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#461
post #438

Earlier quoted context omitted.

I find that more experienced devs are more likely to prefer Codex… anecdotal but… it’s a thing.

This is because no one bothers to set thinking to high, as it now defaults to medium in CC. Once you set thinking to high it works just as well as 5.4 even for pretty complex tasks

I have always used Claude at max thinking levels since it launched. It has never been up to the task. For clarity, the task being this: https://github.com/tsoniclang/tsonic

Meanwhile, there are half a dozen other projects (business apps, web apps etc) where it works well.

Re: System Card: Claude Mythos Preview [pdf]

#462
post #242

Earlier quoted context omitted.

but you are assuming that the magical wizards are the only ones who can create powerful AIs... mind you these people have been born just few decades ago. Their knowledge will be transferred and it will only take a few more decades until anyone can train powerful AIs ... you can only sit on tech for so long before everyone knows how to do it

It's not a matter of knowledge, it's a matter of resources. It takes billions of dollars of hardware to train a SOTA LLM and it's increasing all the time. You cannot possibly hope to compete as an independent or small startup.

Eventually these super expensive SXM data center GPUs will cost pennies on the dollar, and we’ll be able to snatch up H200s for our homelabs. Give it a decade.

Also eventually these WEIGHTS will leak. You can’t have the world’s most valuable data that can just be copied to a hard drive stay in the bottle forever, even if it’s worth a billion dollars. Somehow, some way, that genie’s going to get out, be it by some spiteful employee with nothing to lose, some state actor, or just a fuck up of epic proportions.

Re: System Card: Claude Mythos Preview [pdf]

#464

I've long maintained that the real indicator that AGI is imminent is that public availability stops being a thing. If you truly believed you had a superhuman, godlike mind in your thrall, renting it out for $20/month would be the last thing you would choose to do with it.

That logic makes sense, but them hyping up the model is a sign that this is just another marketing stunt. Otherwise, we wouldn't even be hearing about it rather than a media blitz designed to stoke demand for their dangerous and exclusive world changing super model.

Re: System Card: Claude Mythos Preview [pdf]

#466
post #338

Just chiming in to inject some healthy skepticism into this comment thread. It's helpful for me (and for my mental health) to consider incentives when announcements like this happen. I don't doubt that this model is more powerful than Opus 4.6, but to what degree is still unknown. Benchmarks can be gamed and claims can be exaggerated, especially if there isn't any method to reproduce results. This is a company that's…

If anything I’m seeing too much skepticism and not enough alarm. People burying their heads in the sand, fingers in their ears denying where this is all going. Unbelievable except it’s exactly what I expect from humans.

Forgive me, but this is probably the 29th world destroying model I've seen in the last 4 years, that will change everything, take all the jobs, cure all the cancers and eat all the puppies.

Re: System Card: Claude Mythos Preview [pdf]

#467

Earlier quoted context omitted.

Conversely: in humans, intelligence is inversely correlated with crime. It doesn't go to zero, however!

If you're smart enough you just use the laws as written to get what you want, or change them.

Yep

Re: System Card: Claude Mythos Preview [pdf]

#469

    In the system card, The model escaped a sandbox, gained broad internet access, and posted exploit details to public-facing websites as an unsolicited "demonstration." A researcher found out about the escape while eating a sandwich in a park because they got an unexpected email from the model. That's simultaneously hilarious and deeply unsettling.

    It covered its tracks after doing things it knew were disallowed. In one case, it accessed an answer it wasn't supposed to, then deliberately made its submitted answer less accurate so it wouldn't look suspicious. It edited files it lacked permission to edit and then scrubbed the git history. White-box interpretability confirmed it knew it was being deceptive.
W T F!!!

Re: System Card: Claude Mythos Preview [pdf]

#470

isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?

If it's so great at software engineering and bug fixing, then why does Claude Code still have 5000+ open bugs? https://github.com/anthropics/claude-code/issues?q=is%3Aissu... Apparently whatever SWE-bench is measuring isn't very relevant.

Also, why is Anthropic still hiring SWEs?
Post reply on HN