Live data from Hacker News

System Card: Claude Mythos Preview [pdf]

www-cdn.anthropic.com

421–430 of 687 posts

Re: System Card: Claude Mythos Preview [pdf]

#421

Again, wake me up when it can do laundry.

Time to wake up: π*0.6: two and a half hours of unseen folding laundry (Physical Intelligence) https://www.youtube.com/watch?v=ZpHapIlJnMo

Looks like the first two hours were spent trying to fold the same t-shirt :)

Re: System Card: Claude Mythos Preview [pdf]

#422
post #338

Just chiming in to inject some healthy skepticism into this comment thread. It's helpful for me (and for my mental health) to consider incentives when announcements like this happen. I don't doubt that this model is more powerful than Opus 4.6, but to what degree is still unknown. Benchmarks can be gamed and claims can be exaggerated, especially if there isn't any method to reproduce results. This is a company that's…

I have been thinking that these SWE benchmarks will continue to improve since these companies hire very intelligent software engineers, they can task a multitude of them to solve problems, and then train the model on those answers. Data has always been the core of it all, onward to the next abstraction, I suppose.

I think computational thinking, or basically "how do I solve this problem efficiently" training data is more valuable then feeding in answers. I don't know what these AI models training data consist of, but it would be interesting to see a model trained purely on reasoning, methods, those foundational skills (basic programming? or maybe not) and then give it some benchmarks.

Re: System Card: Claude Mythos Preview [pdf]

#423
post #212

While we still have months to a year or two left, I will once again remind people that it's not too late to change our current trajectory. You are not "anti-progress" to not want this future we are building, as you are not "anti-progress" for not wanting your kids to grow up on smart phones and social media. We should remember that not all technology is net-good for humanity, and this technology in particular poses u…

Just because the path is bad doesn't mean it won't happen.

The other thing you're failing to look at is momentum and majority opinion. When you look at that... nothings going to change, it's like asking an addict to stop using drugs. The end game of AI will play out, that is the most probably outcome. Better to prepare for the end game.

It's similar to global warming. Everyone gets pissed when I say this but the end game for global warming will play out, prevention or mitigation is still possible and not enough people will change their behavior to stop it. Ironically it's everyone thinking like this and the impossibility of stopping everyone from thinking like this that is causing everyone to think and behave like this.

Re: System Card: Claude Mythos Preview [pdf]

#424
post #107

isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?

Anthropic needs to show that its models continually get better. If the model showed minimal to no improvement, it would cause significant damage to their valuation. We have no way of validating any of this, there are no independent researchers that can back any of the assertions made by Anthropic. I don’t doubt they have found interesting security holes, the question is how they actually found them. This System Card…

Well they said theyll be giving the model to select tech companies to use, there soon will be independent users who can comment on its capabilities.

Re: System Card: Claude Mythos Preview [pdf]

#425
post #208

Earlier quoted context omitted.

Alignment “appearing” better as model capabilities increase scares the shit out of me, tbh.

Conversely: in humans, intelligence is inversely correlated with crime. It doesn't go to zero, however!

If you're smart enough you just use the laws as written to get what you want, or change them.

Re: System Card: Claude Mythos Preview [pdf]

#426
post #18

At what point do these companies stop releasing models and just use them to bootstrap AGI for themselves?

Plausibly now. "As we wrote in the Project Glasswing announcement, we do not plan to make Mythos Preview generally available"

I remember when they didn't plan to give LLMs internet access for the same safety reasons.

Re: System Card: Claude Mythos Preview [pdf]

#427

Earlier quoted context omitted.

In what way is AI 2027 coming true? AI 2027 predicted a giant model with the ability to accelerate AI research exponentially. This isn't happening. AI 2027 didn't predict a model with superhuman zero-day finding skills. This is what's happening. Also, I just looked through it again, and they never even predicted when AI would get good at video games. It just went straight from being bad at video games to world domina…

> Early 2026: OpenBrain continues to deploy the iteratively improving Agent-1 internally for AI R&D. Overall, they are making algorithmic progress 50% faster than they would without AI assistants—and more importantly, faster than their competitors. > you could think of Agent-1 as a scatterbrained employee who thrives under careful management According to this document, 1 of the 18 Anthropic staff surveyed even said t…

> According to this document, 1 of the 18 Anthropic staff surveyed even said the model could completely replace an entry level researcher. > > So I'd say we've reached this milestone.

If 1/N=18 are our requirements for statistical significance for world-altering claims, then yeah, I think we can replace all the researchers.

Re: System Card: Claude Mythos Preview [pdf]

#428
post #338

Just chiming in to inject some healthy skepticism into this comment thread. It's helpful for me (and for my mental health) to consider incentives when announcements like this happen. I don't doubt that this model is more powerful than Opus 4.6, but to what degree is still unknown. Benchmarks can be gamed and claims can be exaggerated, especially if there isn't any method to reproduce results. This is a company that's…

What would be the incentive to engage in the tactic when the proof is ultimately in the pudding when the model hits the streets? Who would ultimately benefit from fudging these numbers?

Re: System Card: Claude Mythos Preview [pdf]

#429
post #231

Earlier quoted context omitted.

I've been increasingly "freaking out" since about 3 - 4 years ago and it seems that the pessimistic scenario is materializing. It looks like it will be over for software engineers in a not so distant future. In January 2025 I said that I expect software engineers to be replaced in 2 years (pessimistic) to 5 years (optimistic). Right now I'm guessing 1 to 3 years.

I assure you it will soon become very clear that mass job losses are one of the least concerning side effects of developing the magic "everything that can plausibly been done within the constraints of physics is now possible" machine. We're opening a can of worms which I don't think most people have the imagination to understand the horrors of.

What is the opposite of horrors and why don't we talk about those ever.

Re: System Card: Claude Mythos Preview [pdf]

#430
post #18

At what point do these companies stop releasing models and just use them to bootstrap AGI for themselves?

I think it is naive to think the government (US or China most probably) will just let some random company control something so powerful and dangerous.

I think it is naive to think that artificial super intelligence will be controlled by anyone.

If it is smarter than all humans combined at everything why would any humans collectively control the ai?

All the ants in your backyard still make no decisions vs you

Post reply on HN