Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

131–140 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#131
post #4

> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. A month ago I might have believed this, now I assume that they know they can't handle the demand for the prices they're advertising.

GPT-2, o1, Opus...been here so many times. The reason they do this is because they know it works (and they seem to specifically employ credulous people who are prone to believe AGI is right around the corner). There haven't been significant innovations, the code generated is still not good but the hype cycle has to retrigger. I remember when OpenAI created the first thinking model with o1 and there were all these bre…

> All thinking does is burn output tokens for accuracy

“All that phenomenon X does is make a tradeoff of Y for Z”

It sounds like you’re indignant about it being called thinking, that’s fine, but surely you can realize that the mechanism you’re criticizing actually works really well?

Re: System Card: Claude Mythos Preview [pdf]

#132
post #101

Earlier quoted context omitted.

Well don’t forget we still have competition. Were anthropic to rent seek OpenAI would undercut them. Were OpenAI and anthropic to collude that would be illegal. For anthropic to capture the entire coding agent market and THEN rent seek, these days it’s never been easier to raise $1B and start a competing lab

In practice this doesn't work though, the Mastercard-Visa duopoly is an example, two competing forces doesn't create aggressive enough competition to benefit the consumer. The only hope we have is the Chinese models, but it will always be too expensive to run the full models for yourself.

Chinese competition can always be banned. Example: Chinese electric car competition

Re: System Card: Claude Mythos Preview [pdf]

#134
post #101

Earlier quoted context omitted.

In practice this doesn't work though, the Mastercard-Visa duopoly is an example, two competing forces doesn't create aggressive enough competition to benefit the consumer. The only hope we have is the Chinese models, but it will always be too expensive to run the full models for yourself.

Chinese competition can always be banned. Example: Chinese electric car competition

That's what OP was saying, I think, noting that running them locally won't be a solution.

Re: System Card: Claude Mythos Preview [pdf]

#135

Earlier quoted context omitted.

Just checked my subscription start date for Anthropic. September 2023, I believe before they announced public launch. Sorry kid.

Genuine question - if you don't think the models are improved or that the code is any good, why do you still have a subscription? You must see some value, or are you in a situation where you're required to test / use it, eg to report on it or required by employer? (I would disagree about the code, the benefits seem obvious to me. But I'm still curious why others would disagree, especially after actively using them fo…

The assumption that the other person made was that I would only use it for coding. If you look through my other comments today, I suggest that they are useful for performing repetitive tasks i.e. checking lint on PR, etc. Also, can be used for throwaway code, very useful.

I don't think the issue is with the model, it is with the implication that AGI is just around the corner and that is what is required for AI to be useful...which is not accurate. The more grey area is with agentic coding but my opinion (one that I didn't always hold) is that these workflows are a complete waste of time. The problem is: if all this is true then how does the CTO justify spending $1m/month on Anthropic (I work somewhere where this has happened, OpenAI got the earlier contract then Cursor Teams was added, now they are adding Anthropic...within 72 hours of the rollout, it was pulled back from non-engineering teams). I think companies will ask why they need to pay Anthropic to do a job they were doing without Anthropic six months ago.

Also, the code is bad. This is something that is non-obvious to 95% of people who talk about AI online because they don't work in a team environment or manage legacy applications. If I interview somewhere and they are using agentic workflow, the codebase will be shit and the company will be unable to deliver. At most companies, the average developer is an idiot, giving them AI is like giving a monkey an AK-47 (I also say this as someone of middling competence, I have been the monkey with AK many times). You increase the ability to produce output without improving the ability to produce good output. That is the reality of coding in most jobs.

AI isn't good enough to replace a competent human, it is fast enough to make an incompetent human dangerous.

Re: System Card: Claude Mythos Preview [pdf]

#136
post #129

Earlier quoted context omitted.

I think are fundamental issues with the story that Anthropic is selling. AGI is very close, we will definitely get there, it is also very dangerous...so Anthropic should be the only ones trusted with AGI. If you look at recent changes in Opus behaviour and this model that is, apparently, amazingly powerful but even more unsafe...seems suspect.

> AGI is very close Based on? Or are you just quoting Anthropic here?

My Anthropic rep told me it was just around the corner...you aren't saying he lied to me? Can't believe this, I thought he was my friend.

Re: System Card: Claude Mythos Preview [pdf]

#137
post #28
post #6

Congratulations to the US military, I guess.

Doesn't Anthropic not have that contract anymore, after all that buzz a month or so ago?

The US has invaded two sovereign countries this year to take their oil. I assume taking over a US company for their AI model would be trivial.

Re: System Card: Claude Mythos Preview [pdf]

#139

isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?

I think there's no SOA advance on this one worthy of "freaking out".

Looks like they just built a way larger model, with the same quirks than Claude 4. Seems like a super expensive "Claude 4.7" model.

I have no doubts that Google and OpenAI already done that for internal (or even government) usage.

Post reply on HN