Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

861–870 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#862
post #467

Earlier quoted context omitted.

While I can't speak for everyone in academia, I personally don't feel comfortable in putting my research questions and outputs to a private website, before the idea is at least arxived. Especially as all the Fable/Mythos prompts are said to be human reviewed. So I believe that, at least in the short run, we might be seeing breakthroughs in hard open problems or in low hanging problems which are not that interesting t…

Isn't it showing a problem with an academia? "I don't want to live in a world where someone else makes the world a better place than we do."

I think the GP was suggesting that their reluctance was more about someone taking their idea. I still think it might suggest a problem with academia, but the summary would be closer to

"I don't want to live in a world where someone else follows through with my ideas without giving me credit"

It's still a problem because a lot of academics aren't especially well equipped to follow through with their ideas, which can create information silos that lead to ideas never being implemented. Still, I don't know if this is the biggest fish to fry: you have other silos like IP law and NDAs etc.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#863

Earlier quoted context omitted.

Just the other way I was thinking that if I asked "What does Lamborghini do?" the only correct way to answer is a single sentence "Which Lamborghini are you referring to?". But LLMs will fail at this question: they will tell you about Lamborghini's latest car and mix some history in it. Just try. Which is the wrong answer anyway, because there's at least two major companies called Lamborghini, one making cars, one ma…

This is ... unnecessarily pedantic. Anybody in my social universe who asked me that question would undoubtedly expect "they make cars". If you're picking nits, why not focus on the word "do" and (wrongly) expect an answer like "Lamborghini (either of the two main companies of that name) does not 'do' anything - the companies employ humans who 'do' things. Lamborghini is a legal entity established to allow humans to '…

This is not a nit. This is a real problem in the technology.

It assumes the average and plausible answer token by token.

And this tendency shows in every single field it's applied to.

At the end of the day I want *correct* answers, to the point.

Instead LLMs, no matter if it's version 3500, are bound to producing average results: slop.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#864
post #520

I’ll be very excited to try it out and see the actual improvement in writing style. The denser writing style probably won’t bother me. Anthropic seems to be listening to community complaint on HN about how the writing style is grating. And apparently the solution from Anthropic is to add this block to every conversation!? > Mannered prose substitutes metaphor and flourish for direct statement. Instead of "a parameter…

[dead]

Re: Claude Fable 5.1 and Claude Mythos 5.1

#865

I let it go a few hours on a not trivial but well-known problem, and it felt like it was just a little too plodding and just kind of mucked around a little too much and wasn't aggressive enough about getting stuff done. I asked it to wind it down and finish up and it took another hour and 15 to actually stop and commit without really getting much more done. Not very impressed here, if you can't tell. This new version…

250k tokens is awfully small for a frontier context window, I haven't pushed the frontier models in a long time but i assumed they all had 1M token windows as advertised practically a year ago.

They do, but compaction helps performance. All models degrade in performance the deeper the context. You can disable auto-compaction, but most of the time you want to /compact, at least, and likely /clear between discrete units of work.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#866
I asked Fable 5.1 (Extra) to review the complaints on hackernews, it scanned everything and then came up with some benign (but scary sounding) prompts that used to get blocked on 5.0. I asked it to give results and it happily answered the bio/cyber prompts - hardened a Dockerfile, did seccomp, found a command injection in its own snippet, etc. Nothing was blocked.

Then i went and pasted those exact prompts it generated into a fresh chat, 5.1 on max effort.

Immediately got blocked and sent to Opus 4.8 fallback.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#867
post #243

I've recently been running these agent sessions on more and more long running tasks because these latest models can do a REALLY good job on big chunks of work, and i've been watching them way less. It's starting to occur to me the importance of alignment is a today problem, it's not a tomorrow problem. In the past I watched and saw everything the model did, not a lot got past me. Today it does A TON of work while i'm…

what do you do while agents work?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#868

Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke. What they have done: * Nerfed Fable, as many of noted it's useless * Leverage Mythos as a marketing strategy, claiming its too good to release * Removed thought traces, one of the only useful things to make sure your prompts are working correctly…

> Push a bunch of EU Overregulation onto the rest of the world with text watermarking, That's not part of the EU regulations. You only need to say that it is created by AI, and then only under certain conditions.

That is simply not true. You need to go read that again, if you ever read it at all before correcting somebody about it

https://artificialintelligenceact.eu/article/50/

"Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated. Providers shall ensure their technical solutions are effective, interoperable, robust and reliable as far as this is technically feasible, taking into account the specificities and limitations of various types of content, the costs of implementation and the generally acknowledged state of the art, as may be reflected in relevant technical standards."

Eg... Watermarking...

Re: Claude Fable 5.1 and Claude Mythos 5.1

#869

If I am reading this right, Fable 5 was worse than Opus 5 in almost every category, while consuming twice the tokens? The things you learn every day...

One day HN will learn that benchmarks aren't everything! Fable 5 was way better than Opus 5 for anyone that used it.

I’m sure they’re not, and it felt the same to me too.

The issue I had with Fable 5 and will probably carry to Fable 5.1 is that I hit the safeguards too often.

That being said, even if the benchmarks are only part of the story, these ones paint a pretty compelling one, when comparing Fable 5.1 to 5

Re: Claude Fable 5.1 and Claude Mythos 5.1

#870
post #405

I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets. I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited. I urge Anthropic to get be…

> I am becoming dependent on AI to make a living IMO, if you depend on AI to make a living, I'd invest in hardware for local inference, and learn on how to effectively make a living using AI inference you control, on hardware you control. Sure, economically speaking it's way cheaper to use one of these heavily subsidised services (for now), and their models are faster and more capable, but if your livelihood depends…

Except there's a huge gulf of self-hosting and using API hosts - no way you can reach the economics of a shared host. Privacy is a problem but you can chose who you host with and where it's hosted (which jurisdiction).

When privacy/compliance really starts to matter it's up to the client/business to provide you with tooling - you're not running that on your own hardware anyway.

So the local AI for individuals is just a hobby/gimmick at this point not a rational decision. Self-hosting for business is a different story.

Post reply on HN