Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

561–570 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#561
post #64
post #18

At what point do these companies stop releasing models and just use them to bootstrap AGI for themselves?

Fictional timeline that holds up pretty well so far: https://ai-2027.com/

"So far" is two entries: "AI companies build bigger datacenters" and "AI is being used for AI research with modest success".

Re: System Card: Claude Mythos Preview [pdf]

#562
post #121

Earlier quoted context omitted.

Can LLMs be AGI at all?

What can a SOTA LLM not answer that the average person can? It's already more intelligent than any polymath that ever existed, it just lacks motivation and agency.

And has ADHD, but yeah, I'm fairly convinced that AGI is already here.

Re: System Card: Claude Mythos Preview [pdf]

#563
post #444

Earlier quoted context omitted.

> A System „Card“ spanning 244 pages. Probably because they asked Claude to write it.

I read the entire thing fwiw (pseudo-retired life helps with time here). It looks like it was a collaborative effort across multiple teams, where each team (research, security, psycology, etc etc etc) were all submitting ~10 pages or so. It doesn't feel like slop.

AI writing has stopped feeling like slop around Opus 4.5, though.

Re: System Card: Claude Mythos Preview [pdf]

#564
post #471

Earlier quoted context omitted.

I've been increasingly "freaking out" since about 3 - 4 years ago and it seems that the pessimistic scenario is materializing. It looks like it will be over for software engineers in a not so distant future. In January 2025 I said that I expect software engineers to be replaced in 2 years (pessimistic) to 5 years (optimistic). Right now I'm guessing 1 to 3 years.

> I've been increasingly "freaking out" since about 3 - 4 years ago and it seems that the pessimistic scenario is materializing. It looks like it will be over for software engineers in a not so distant future. In January 2025 I said that I expect software engineers to be replaced in 2 years (pessimistic) to 5 years (optimistic). Right now I'm guessing 1 to 3 years. Tell me how this will replace Jira, planning, convin…

Have you never filed JIRA tickets, planned, or debated viability with an AI? Which part of those are you finding that an AI absolutely cannot do better than the average developer?

Re: System Card: Claude Mythos Preview [pdf]

#565
post #371

Earlier quoted context omitted.

What sources would you even be looking for? I think you're asking the wrong question. It's not like I'm arguing a scientific theory which can be backed by data and experimentation. I can only provide you reasoning for why I believe what I believe. Firstly, I'd propose that all technological advances are a product of time and intelligence, and that given unlimited time and intelligence, the discovery and application o…

You are supposing it's possible to know that much about some things that maybe are not knowledgeable to us, even with these tools. Life is extremely complex, more than it's typically assumed by engineering-minded people. Let's be humble here and acknowledge it.

Life might be complex, but it isn't unknowable. Claiming life is unknowable isn't being humble, it's being naive.

Re: System Card: Claude Mythos Preview [pdf]

#566

Earlier quoted context omitted.

Good catch. If it's "too slow" even when ran in a state-of-the-art datacenter environment, this "Mythos" model is most closely comparable to the "Deep Research" modes for GPT and Gemini, which Claude formerly lacked any direct equivalent for.

I don't think that's what's being hinted at. The system card seems to say that the model is both token efficient and slow in practice. Deep research modes generally work by having many subagents/large token spend. So this more likely the fact that each token just takes longer to produce, which would be because the model is simply much larger. By epoch AIs datacenter tracking methods, anthropic has had access to the l…

"Slow and token-efficient" could be achieved quite trivially by taking an existing large MoE model and increasing the amount of active experts per layer, thus decreasing sparsity. The broader point is that to end users, Mythos behaves just like Deep Research: having it be "more token efficient" compared to running swarms of subagents is not something that impacts them directly.

Re: System Card: Claude Mythos Preview [pdf]

#567

Earlier quoted context omitted.

This is the notebook filled with exposition you find in post apocalyptic videogames.

Everything they built. Imperfect. So easy to take control.

They think that they are safe. They are not.

Re: System Card: Claude Mythos Preview [pdf]

#568
post #266
post #223

Earlier quoted context omitted.

My understanding is GPT 6 works via synaptic space reasoning... which I find terrifying. I hope if true, OpenAI does some safety testing on that, beyond what they normally do.

From the recent New Yorker piece on Sam: “My vibes don’t match a lot of the traditional A.I.-safety stuff,” Altman said. He insisted that he continued to prioritize these matters, but when pressed for specifics he was vague: “We still will run safety projects, or at least safety-adjacent projects.” When we asked to interview researchers at the company who were working on existential safety—the kinds of issues that co…

Why are these people always like this.

Re: System Card: Claude Mythos Preview [pdf]

#569
post #410

Earlier quoted context omitted.

The day I start freaking out about my job is the day when my non-engineer friend turned vibe coder understands how, or why the thing that AI wrote works. Or why something doesn't work exactly the way he envisioned and what does it take to get it there. If it can replace SWEs, then there's no reason why it can't replace say, a lawyer, or any other job for that matter. If it can't, then SWE is fine. If it can - well, w…

> If it can replace SWEs, then there's no reason why it can't replace say, a lawyer SWE is unique in that for part of the job it's possible to set up automated verification for correct output - so you can train a model to be better at it. I don't think that exists in law or even most other work.

What is the automated verification of correct output and who defines that?

But before verification, what IS correct output?

I understand SWE process is unique in that there are some automations that verify some inputs and outputs, but this reasoning falls into the same fallacies that we've had before AI era. First one that comes to mind is that 100% code coverage in tests means that software is perfect.

Post reply on HN