Live data from Hacker News

Claude Opus 4.7 Model Card

anthropic.com

51–60 of 88 posts

Re: Claude Opus 4.7 Model Card

#51
post #12

Dumb question but why are chemical weapons always addressed as a risk with llms? Is the idea that they contain how to make chemical weapons or that they would guide someone on how? Would there not already be websites that contain that information? How is an llm different, i guess, from some sort of anarchist cookbook thing.

It’s marketing, Fear is one of the most effective marketing tools. That and purpose of government attention

Re: Claude Opus 4.7 Model Card

#52
post #49

Earlier quoted context omitted.

That’s not quite true. Take a look at all the billionaires destroying society. Being evil is the surest way to get to get rich. In fact it’s the only way to amass that level of capital: there’s no ethical billionaire.

This feels like a wild overgeneralization. People can become rich without resorting to evil methods, especially now with global markets and software. Case in point: Minecraft was wildly successful, and now Notch is a billionaire.

Eeeeh not the best example maybe?

Re: Claude Opus 4.7 Model Card

#53
post #6

This reads more like an advertisement for Mythos, on the first glance

I never understand these critiques. If something is useful and you’re selling it, does that mean any technical document describing its usefulness becomes marketing?

I guess maybe, but then do those documents lose value as technical documents? Not necessarily at all, so I don’t see the point. How are you supposed to describe a useful technical thing to users?

Re: Claude Opus 4.7 Model Card

#54

So Opus 4.7 is measurably worse at long-context retrieval compared to Opus 4.6. Opus 4.6 scores 91.9% and Opus 4.7 scores 59.2%. At least they're transparent about the model degradation. They traded long-context retrieval for better software engineering and math scores.

A year ago it felt like SoTA model developers were not improving so much as moving the dirt around. Maybe we’re in another such rut.

Re: Claude Opus 4.7 Model Card

#56

Earlier quoted context omitted.

"Smart people have economic opportunities that align them away from being evil" For some definition of evil, some of the time, ok. But as economic opportunities compound (looking at the behavior of the ultra-rich), it seems there's at least strong correlation in the other direction, if not full-on "root of all evil" causation.

Sure, but that’s not “slaughter a stadium of people with drones” evil or “poison the water supply” evil or “take out unprotected electrical substations” evil. So much infrastructure is very soft because the evil people aren’t smart enough to conceive of or conduct an attack.

I think you might find that, if you reconsider who the 'evil' people are, you might find that we're already doing that sort of thing.

Re: Claude Opus 4.7 Model Card

#57
post #49

Earlier quoted context omitted.

This feels like a wild overgeneralization. People can become rich without resorting to evil methods, especially now with global markets and software. Case in point: Minecraft was wildly successful, and now Notch is a billionaire.

Eeeeh not the best example maybe?

Pre-wealth, Notch was friendly, kind, and downright jolly! Even as he started to accumulate wealth, he was donating huge sums of money to various indie games. Whenever a Humble Bundle dropped he would top the leaderboard for the amount he paid for the games. Things took a major turn for the worse after the acquisition and after he left Mojang. That's when he ran out of purpose and turned to drugs and conspiracy theories.

Re: Claude Opus 4.7 Model Card

#58
post #12

Dumb question but why are chemical weapons always addressed as a risk with llms? Is the idea that they contain how to make chemical weapons or that they would guide someone on how? Would there not already be websites that contain that information? How is an llm different, i guess, from some sort of anarchist cookbook thing.

They contain broad overviews(throw some disease-causing bacteria in a sort of rainbow arrangement of increasingly more effective antibiotics, you'll usually get something that's at least very deadly even if it doesn't have pandemic potential) but executing in a real lab takes a ton of trial and error to figure out the details. The issue is that the details ~all exist somewhere in the training dataset already, discovered and documented over the course of unrelated, benign biology research. Ability to quickly and accurately search over that corpus translates to large speedups in the physical development process.

Re: Claude Opus 4.7 Model Card

#59
post #53
post #6

This reads more like an advertisement for Mythos, on the first glance

I never understand these critiques. If something is useful and you’re selling it, does that mean any technical document describing its usefulness becomes marketing? I guess maybe, but then do those documents lose value as technical documents? Not necessarily at all, so I don’t see the point. How are you supposed to describe a useful technical thing to users?

This is supposedly the Opus 4.7 model card. It's okay for it to be marketing for Opus 4.7 and describe what it can do, and even okay for it to talk about what it does better than the last generation. GP was saying it sounds like marketing for Mythos (a different and unreleased model). I don't want the Opus 4.7 model card to be advertising for something else.

For context, the word "Mythos" appears 331 times in a 221 page document. "Opus 4.6" appears 240 times, so a reference to a model that nobody has really used happens more often than the reference to the last generation model.

Re: Claude Opus 4.7 Model Card

#60

So Opus 4.7 is measurably worse at long-context retrieval compared to Opus 4.6. Opus 4.6 scores 91.9% and Opus 4.7 scores 59.2%. At least they're transparent about the model degradation. They traded long-context retrieval for better software engineering and math scores.

Be brief. No one wants AI boyfriend users who drone on & on about their day.
Post reply on HN