Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

261–270 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#261

> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.

I sometimes have that feeling too, then ask another LLM to do a code and vulnerability review and OMG: rookie mistakes, over complications and security gaps even a 1st year student would not make regularly. So.. one more year of untreated bipolar AI psychosis I guess..

I think these kinds of comments really need to say which LLM that is. There's an enormous difference in skill between the frontier ones and say the Google search AI.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#262

Instead of a new model that's going to have unreasonably shallow usage limits, I wish they would: 1) address the claude 20x plan usage being only 6-7x the ceiling of the claude pro plan 2) either fix opus 5, make it completely free, or delete it entirely

> either fix opus 5, make it completely free, or delete it entirely

They should pay for us for using it!

Re: Claude Fable 5.1 and Claude Mythos 5.1

#263

> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.

We will have more bugs. Even the best models with the best software engineers will produce bugs. There are two reasons : first the pressure to produce more and second LLMs will always produce slop

Re: Claude Fable 5.1 and Claude Mythos 5.1

#264

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

Serious question: Do you suffer internally from too much slop being submitted? How do you counter that?

Context:

If you want or not, many engineers will eventually end up sending ai slop to your PR or maybe even skip and trigger CI/CD.

Many company owners, OSS maintainers and projects suffer from slop-code being submitted in high-frequency.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#265

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

Nice, I'm looking forward to the improved writing on the majority of articles posted here.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#267

> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.

I think we'll have lots of bugs. They'll just be found and closed way sooner. You'll have an agent that watchs for issues, then opens a PR fixing it.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#268

Earlier quoted context omitted.

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

> They're packing lots of signal into fewer words There's a huge difference between the kind of prose you see in final output vs CoT windows. The final output is very much not what I'd call "packing lots of signal into fewer words" (aside perhaps from "Claude-isms" being easy enough to scan for if for some reason you actually wanted to scan for them, which other agents might want to for all I know); and if agents are…

I don't know, I just pulled up the status for an active session and here's what it said:

  One thing I found before dispatching, and filed as Q0579. The halt told you C6
  was all that was left in the unit. That was true of the step's criteria and
  false of the unit's acceptance, which reads "exits 0 AND witnessed red" — two
  conjuncts. The witness half holds; the exits-0 half does not, because hello's
  G7 currently reads DIFFER 554/51340. I re-derived that from the gate map
  rather than trusting the prior step's report. So satisfying C6 does not by
  itself finish this unit, and I've filed that so attempt 1's success can't
  quietly be read as the unit's.
It's not exactly plain language.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#269
post #35

“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.” I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better. I don’t think Anthropic realizes that…

I've developed a habit of adding into my prompts "please keep your response concise and succinct" or "I'm trying to cram, please only provide the minimum level of technical detail necessary to understand this topic" I find it helps immensely but it'd be nice if I didn't have to do that.

add to your system prompt?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#270

> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.

The marketing here trick is, if they spent the same money on humans they'd have found it years ago.

Instead, the lurking variable here is new budget was added. With the new budget, they added a new tool, and the bug was located.

The difference here was budget.

Post reply on HN