Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

191–200 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#191
I've long maintained that the real indicator that AGI is imminent is that public availability stops being a thing. If you truly believed you had a superhuman, godlike mind in your thrall, renting it out for $20/month would be the last thing you would choose to do with it.

Re: System Card: Claude Mythos Preview [pdf]

#192
I wonder what the relationship is between a model's capability and the personality it develops.

Page 202:

> In interactions with subagents, internal users sometimes observed that Mythos Preview appeared “disrespectful” when assigning tasks. It showed some tendency to use commands that could be read as “shouty” or dismissive, and in some cases appeared to underestimate subagent intelligence by overexplaining trivial things while also underexplaining necessary context.

Page 207:

> Emoji frequency spans more than two orders of magnitude across models: Opus 4.1 averages 1,306 emoji per conversation, while Mythos Preview averages 37, and Opus 4.5 averages 0.2. Models have their own distinctive sets of emojis: the cosmic set () favored by older models like Sonnet 4 and Opus 4 and 4.1, the functional set () used by Opus 4.5 and 4.6 and Claude Sonnet 4.5, and Mythos Preview's “nature” set ().

Re: System Card: Claude Mythos Preview [pdf]

#193

Earlier quoted context omitted.

"some model I don't get to use is much better at benchmarks" pick one or more: comically huge model, test time scaling at 10e12W, benchmark overfit

So... you're not excited because it might take a few months before we can use it or something? I don't get your comment.

Whether you're excited depends on what do you do for living and how close you are to financial independence.

Re: System Card: Claude Mythos Preview [pdf]

#194

Earlier quoted context omitted.

Describing providing a highly valuable service for money as `rent seeking` is pretty wild.

My housing is pretty valuable. I pay rent. Which timeline are you in?

Rent seeking refers to https://en.wikipedia.org/wiki/Rent-seeking

Re: System Card: Claude Mythos Preview [pdf]

#195

Earlier quoted context omitted.

Man... It's hard after seeing this to not be worried about the future of SWE If AI really is bench marking this well -> just sell it as a complete replacement which you can charge for some insane premium, just has to cost less than the employees... I was worried before, but this is truly the darkest timeline if this is really what these companies are going for.

Of course it's what they're going for. If they could do it they'd replace all human labor - unfortunately it's looking like SWE might be the easiest of the bunch. The weirdest thing to me is how many working SWEs are actively supporting them in the mission.

Enthusiastically supporting them. It’s quite depressing to watch over the last few years. It’s not like they’re being coy about their aim…

Re: System Card: Claude Mythos Preview [pdf]

#196
post #130

Earlier quoted context omitted.

So did GPT-4. https://arxiv.org/html/2402.06664v1 Like think carefully about this. Did they discover AGI? Or did a bunch of investors make a leveraged bet on them "discovering AGI" so they're doing absolutely anything they can to make it seem like this time it's brand new and different. If we're to believe Anthropic on these claims, we also have to just take it on faith, with absolutely no evidence, that they've made…

I don't see the problem here. How would you have handled it differently? If you released this model as such without any safety concern, the vulnerabilities might be found by bad actors and used for wrong things. What do you find surprising here?

Vulnerabilities were found, probably a few by bad actors, when GPT4 was released. Every vulnerability found now is probably found with AI assistance at the very least. Should they have never released GPT4? Should we have believed claims that GPT4 was too dangerous for mere mortals to access? I believe openAI was making similar claims about how GPT4 was a step function and going to change white collar work forever when that model was released.

The point is that this whole "the model is too powerful" schtick is a bunch of smoke and mirrors. It serves the valuation.

Re: System Card: Claude Mythos Preview [pdf]

#197
post #67

Earlier quoted context omitted.

Anyone who has used Opus recently can verify that their current model does all of these things quite competently.

That has also been my experience. And if Mythos is even worse, unless you have a significantly awesome harness, sounds like pretty unusable if you don't want to risk those problems.

Human in the loop is the best way to go. You'll still be way faster than without the agent, and there is no risk of it going haywire unless you turn off your brain!

Re: System Card: Claude Mythos Preview [pdf]

#199

A System „Card“ spanning 244 pages. Quite a stretch of the original word meaning.

In corporate circles there is an allergy to use "request" ("ask" is used as a noun) and "lesson" ("learning" has been invented for the same role).

I guess now anything that sounds related to school will be banned so "book" is on its way out.

Re: System Card: Claude Mythos Preview [pdf]

#200

Earlier quoted context omitted.

Describing providing a highly valuable service for money as `rent seeking` is pretty wild.

My housing is pretty valuable. I pay rent. Which timeline are you in?

Actually you're saying similar things:

Rent-seeking of old was a ground rent, monies paid for the land without considering the building that was on it.

Residential rents today often have implied warrants because of modern law, so your landlord is essentially selling you a service at a particular location.

Post reply on HN