Claude Fable 5.1 and Claude Mythos 5.1
861–870 of 1001 posts
Re: Claude Fable 5.1 and Claude Mythos 5.1
#862Earlier quoted context omitted.
While I can't speak for everyone in academia, I personally don't feel comfortable in putting my research questions and outputs to a private website, before the idea is at least arxived. Especially as all the Fable/Mythos prompts are said to be human reviewed. So I believe that, at least in the short run, we might be seeing breakthroughs in hard open problems or in low hanging problems which are not that interesting t…
Isn't it showing a problem with an academia? "I don't want to live in a world where someone else makes the world a better place than we do."
"I don't want to live in a world where someone else follows through with my ideas without giving me credit"
It's still a problem because a lot of academics aren't especially well equipped to follow through with their ideas, which can create information silos that lead to ideas never being implemented. Still, I don't know if this is the biggest fish to fry: you have other silos like IP law and NDAs etc.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#863Earlier quoted context omitted.
Just the other way I was thinking that if I asked "What does Lamborghini do?" the only correct way to answer is a single sentence "Which Lamborghini are you referring to?". But LLMs will fail at this question: they will tell you about Lamborghini's latest car and mix some history in it. Just try. Which is the wrong answer anyway, because there's at least two major companies called Lamborghini, one making cars, one ma…
This is ... unnecessarily pedantic. Anybody in my social universe who asked me that question would undoubtedly expect "they make cars". If you're picking nits, why not focus on the word "do" and (wrongly) expect an answer like "Lamborghini (either of the two main companies of that name) does not 'do' anything - the companies employ humans who 'do' things. Lamborghini is a legal entity established to allow humans to '…
It assumes the average and plausible answer token by token.
And this tendency shows in every single field it's applied to.
At the end of the day I want *correct* answers, to the point.
Instead LLMs, no matter if it's version 3500, are bound to producing average results: slop.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#864I’ll be very excited to try it out and see the actual improvement in writing style. The denser writing style probably won’t bother me. Anthropic seems to be listening to community complaint on HN about how the writing style is grating. And apparently the solution from Anthropic is to add this block to every conversation!? > Mannered prose substitutes metaphor and flourish for direct statement. Instead of "a parameter…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#865I let it go a few hours on a not trivial but well-known problem, and it felt like it was just a little too plodding and just kind of mucked around a little too much and wasn't aggressive enough about getting stuff done. I asked it to wind it down and finish up and it took another hour and 15 to actually stop and commit without really getting much more done. Not very impressed here, if you can't tell. This new version…
250k tokens is awfully small for a frontier context window, I haven't pushed the frontier models in a long time but i assumed they all had 1M token windows as advertised practically a year ago.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#866Then i went and pasted those exact prompts it generated into a fresh chat, 5.1 on max effort.
Immediately got blocked and sent to Opus 4.8 fallback.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#867I've recently been running these agent sessions on more and more long running tasks because these latest models can do a REALLY good job on big chunks of work, and i've been watching them way less. It's starting to occur to me the importance of alignment is a today problem, it's not a tomorrow problem. In the past I watched and saw everything the model did, not a lot got past me. Today it does A TON of work while i'm…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#868Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke. What they have done: * Nerfed Fable, as many of noted it's useless * Leverage Mythos as a marketing strategy, claiming its too good to release * Removed thought traces, one of the only useful things to make sure your prompts are working correctly…
> Push a bunch of EU Overregulation onto the rest of the world with text watermarking, That's not part of the EU regulations. You only need to say that it is created by AI, and then only under certain conditions.
https://artificialintelligenceact.eu/article/50/
"Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated. Providers shall ensure their technical solutions are effective, interoperable, robust and reliable as far as this is technically feasible, taking into account the specificities and limitations of various types of content, the costs of implementation and the generally acknowledged state of the art, as may be reflected in relevant technical standards."
Eg... Watermarking...
Re: Claude Fable 5.1 and Claude Mythos 5.1
#869If I am reading this right, Fable 5 was worse than Opus 5 in almost every category, while consuming twice the tokens? The things you learn every day...
One day HN will learn that benchmarks aren't everything! Fable 5 was way better than Opus 5 for anyone that used it.
The issue I had with Fable 5 and will probably carry to Fable 5.1 is that I hit the safeguards too often.
That being said, even if the benchmarks are only part of the story, these ones paint a pretty compelling one, when comparing Fable 5.1 to 5
Re: Claude Fable 5.1 and Claude Mythos 5.1
#870I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets. I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited. I urge Anthropic to get be…
> I am becoming dependent on AI to make a living IMO, if you depend on AI to make a living, I'd invest in hardware for local inference, and learn on how to effectively make a living using AI inference you control, on hardware you control. Sure, economically speaking it's way cheaper to use one of these heavily subsidised services (for now), and their models are faster and more capable, but if your livelihood depends…
When privacy/compliance really starts to matter it's up to the client/business to provide you with tooling - you're not running that on your own hardware anyway.
So the local AI for individuals is just a hobby/gimmick at this point not a rational decision. Self-hosting for business is a different story.