Again, wake me up when it can do laundry.
Time to wake up: π*0.6: two and a half hours of unseen folding laundry (Physical Intelligence) https://www.youtube.com/watch?v=ZpHapIlJnMo
System Card: Claude Mythos Preview [pdf]
421–430 of 687 posts
Re: System Card: Claude Mythos Preview [pdf]
#422Just chiming in to inject some healthy skepticism into this comment thread. It's helpful for me (and for my mental health) to consider incentives when announcements like this happen. I don't doubt that this model is more powerful than Opus 4.6, but to what degree is still unknown. Benchmarks can be gamed and claims can be exaggerated, especially if there isn't any method to reproduce results. This is a company that's…
I have been thinking that these SWE benchmarks will continue to improve since these companies hire very intelligent software engineers, they can task a multitude of them to solve problems, and then train the model on those answers. Data has always been the core of it all, onward to the next abstraction, I suppose.
Re: System Card: Claude Mythos Preview [pdf]
#423While we still have months to a year or two left, I will once again remind people that it's not too late to change our current trajectory. You are not "anti-progress" to not want this future we are building, as you are not "anti-progress" for not wanting your kids to grow up on smart phones and social media. We should remember that not all technology is net-good for humanity, and this technology in particular poses u…
The other thing you're failing to look at is momentum and majority opinion. When you look at that... nothings going to change, it's like asking an addict to stop using drugs. The end game of AI will play out, that is the most probably outcome. Better to prepare for the end game.
It's similar to global warming. Everyone gets pissed when I say this but the end game for global warming will play out, prevention or mitigation is still possible and not enough people will change their behavior to stop it. Ironically it's everyone thinking like this and the impossibility of stopping everyone from thinking like this that is causing everyone to think and behave like this.
Re: System Card: Claude Mythos Preview [pdf]
#424isn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?
Anthropic needs to show that its models continually get better. If the model showed minimal to no improvement, it would cause significant damage to their valuation. We have no way of validating any of this, there are no independent researchers that can back any of the assertions made by Anthropic. I don’t doubt they have found interesting security holes, the question is how they actually found them. This System Card…
Re: System Card: Claude Mythos Preview [pdf]
#425Earlier quoted context omitted.
Alignment “appearing” better as model capabilities increase scares the shit out of me, tbh.
Conversely: in humans, intelligence is inversely correlated with crime. It doesn't go to zero, however!
Re: System Card: Claude Mythos Preview [pdf]
#426At what point do these companies stop releasing models and just use them to bootstrap AGI for themselves?
Plausibly now. "As we wrote in the Project Glasswing announcement, we do not plan to make Mythos Preview generally available"
Re: System Card: Claude Mythos Preview [pdf]
#427Earlier quoted context omitted.
In what way is AI 2027 coming true? AI 2027 predicted a giant model with the ability to accelerate AI research exponentially. This isn't happening. AI 2027 didn't predict a model with superhuman zero-day finding skills. This is what's happening. Also, I just looked through it again, and they never even predicted when AI would get good at video games. It just went straight from being bad at video games to world domina…
> Early 2026: OpenBrain continues to deploy the iteratively improving Agent-1 internally for AI R&D. Overall, they are making algorithmic progress 50% faster than they would without AI assistants—and more importantly, faster than their competitors. > you could think of Agent-1 as a scatterbrained employee who thrives under careful management According to this document, 1 of the 18 Anthropic staff surveyed even said t…
If 1/N=18 are our requirements for statistical significance for world-altering claims, then yeah, I think we can replace all the researchers.
Re: System Card: Claude Mythos Preview [pdf]
#428Just chiming in to inject some healthy skepticism into this comment thread. It's helpful for me (and for my mental health) to consider incentives when announcements like this happen. I don't doubt that this model is more powerful than Opus 4.6, but to what degree is still unknown. Benchmarks can be gamed and claims can be exaggerated, especially if there isn't any method to reproduce results. This is a company that's…
Re: System Card: Claude Mythos Preview [pdf]
#429Earlier quoted context omitted.
I've been increasingly "freaking out" since about 3 - 4 years ago and it seems that the pessimistic scenario is materializing. It looks like it will be over for software engineers in a not so distant future. In January 2025 I said that I expect software engineers to be replaced in 2 years (pessimistic) to 5 years (optimistic). Right now I'm guessing 1 to 3 years.
I assure you it will soon become very clear that mass job losses are one of the least concerning side effects of developing the magic "everything that can plausibly been done within the constraints of physics is now possible" machine. We're opening a can of worms which I don't think most people have the imagination to understand the horrors of.
Re: System Card: Claude Mythos Preview [pdf]
#430At what point do these companies stop releasing models and just use them to bootstrap AGI for themselves?
I think it is naive to think the government (US or China most probably) will just let some random company control something so powerful and dangerous.
If it is smarter than all humans combined at everything why would any humans collectively control the ai?
All the ants in your backyard still make no decisions vs you