Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

311–320 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#311

I've long maintained that the real indicator that AGI is imminent is that public availability stops being a thing. If you truly believed you had a superhuman, godlike mind in your thrall, renting it out for $20/month would be the last thing you would choose to do with it.

You would if there was one other company with a just as capable god like AI. You’d undercut them by 500 which would make them undercut you. Do that a couple of times and boom. 20 dollars.

That's still assuming that they're competing as consumer tools, rather than competing to discover the next miracle drug or trading algorithm or whatever. The idea is that there'd more profitable uses for a super-intelligent computer, even if there were more than one.

Re: System Card: Claude Mythos Preview [pdf]

#312

Earlier quoted context omitted.

Are these fair comparisons? It seems like mythos is going to be like a 5.4 ultra or Gemini Deepthink tier model, where access is limited and token usage per query is totally off the charts.

There are a few hints in the doc around this > Importantly, we find that when used in an interactive, synchronous, “hands-on-keyboard” pattern, the benefits of the model were less clear. When used in this fashion, some users perceived Mythos Preview as too slow and did not realize as much value. Autonomous, long-running agent harnesses better elicited the model’s coding capabilities. (p201) ^^ From the surrounding co…

Good catch. If it's "too slow" even when ran in a state-of-the-art datacenter environment, this "Mythos" model is most closely comparable to the "Deep Research" modes for GPT and Gemini, which Claude formerly lacked any direct equivalent for.

Re: System Card: Claude Mythos Preview [pdf]

#313

> Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin. We believe that it does not have any significant coherent misaligned goals, and its character traits in typical conversations closely follow the goals we laid out in our constitution. Even so, we believe that it likely poses the greatest alignment-related risk of any…

There is some unintentional good marketing here -- the model is so good its dangerous.

Reminds me of the book 48 Laws of Power -- so good its banned from prisons.

Re: System Card: Claude Mythos Preview [pdf]

#314

Earlier quoted context omitted.

You have to recoup your training costs though? But I’m sure you would have better option than renting it to the general public if you indeed have a perfected AI

If you truly have an artificial superhuman mind, you don't need to rent it out to profit from it. You can skip to the chase and just have it run businesses itself, instead of renting it to human entrepreneur middlemen.

Running businesses and dealing with customers can be a major pain. There’s a lot of soft work in any business on top of the technical work.

Why bother with all that when you can simply charge an extortionate rate and customers will pay it anyway because it’s still profitable?

Re: System Card: Claude Mythos Preview [pdf]

#315

So what changed? They are surely not getting new data to train with, what is the change in architecture that caused this? Do we not know anything about this model? My fear is Anthropic cannot be the only one that achieved it, OpenAI, Gemini and even the Chinese companies see this and probably achieved it too. At which point not releasing will become moot.

Assuming it's #1 a bigger model (given that it is slower), I'm sure there are a variety of improvements but basically they probably mostly come down to: Scaling keeps working. Are there fundamental improvements though? I don't see signs of it.

Re: System Card: Claude Mythos Preview [pdf]

#316
post #266
post #223

Earlier quoted context omitted.

My understanding is GPT 6 works via synaptic space reasoning... which I find terrifying. I hope if true, OpenAI does some safety testing on that, beyond what they normally do.

From the recent New Yorker piece on Sam: “My vibes don’t match a lot of the traditional A.I.-safety stuff,” Altman said. He insisted that he continued to prioritize these matters, but when pressed for specifics he was vague: “We still will run safety projects, or at least safety-adjacent projects.” When we asked to interview researchers at the company who were working on existential safety—the kinds of issues that co…

Amusing! Even if they believe that, they should know the company communicated the opposite earlier.

Re: System Card: Claude Mythos Preview [pdf]

#317
post #168

Earlier quoted context omitted.

"We want to see risks in the models, so no matter how good the performance and alignment, we’ll see risks, results and reality be damned."

i mean, to be fair, these are professional researchers. i'm very inclined to trust them on the various ways that models can subtly go wrong, in long-term scenarios for example, consider using models to write email -- is it a misalignment problem if the model is just too good at writing marketing emails?? or too good at getting people to pay a spammy company? another hot use case: biohacking. if a model is used to do…

"for example, consider using models to write email -- is it a misalignment problem if the model is just too good at writing marketing emails?? or too good at getting people to pay a spammy company?"

But who gets to be the judge of that kind of "misalignment"? giant tech companies?

Re: System Card: Claude Mythos Preview [pdf]

#318
post #242

Earlier quoted context omitted.

It's not a matter of knowledge, it's a matter of resources. It takes billions of dollars of hardware to train a SOTA LLM and it's increasing all the time. You cannot possibly hope to compete as an independent or small startup.

Presumably, the hardware to run this level of model will be democratized within the timeframe of the parent comment.

See https://amppublic.com and Stanford CS153, https://www.youtube.com/watch?v=mZqh7emiz9Q

Re: System Card: Claude Mythos Preview [pdf]

#319

> Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin. We believe that it does not have any significant coherent misaligned goals, and its character traits in typical conversations closely follow the goals we laid out in our constitution. Even so, we believe that it likely poses the greatest alignment-related risk of any…

There is some unintentional good marketing here -- the model is so good its dangerous. Reminds me of the book 48 Laws of Power -- so good its banned from prisons.

Unintentional? This sort of marketing has been both Antrhopic's and OpenAI's MO for years...

Re: System Card: Claude Mythos Preview [pdf]

#320

So what changed? They are surely not getting new data to train with, what is the change in architecture that caused this? Do we not know anything about this model? My fear is Anthropic cannot be the only one that achieved it, OpenAI, Gemini and even the Chinese companies see this and probably achieved it too. At which point not releasing will become moot.

Well the important thing is they have a lot more data of people actually using their models. They have read billions more lines of private repos and implemented millions of patches, all of which is feeding into the newer models.

More importantly it understand what behaviour people tend to appreciate and what changes are more likely to get approved. This real world usage data is invaluable.

Post reply on HN