Live data from Hacker News

Claude Fable 5: mid-tier results on coding tasks

endorlabs.com

241–250 of 271 posts

Re: Claude Fable 5: mid-tier results on coding tasks

#241
post #227

Earlier quoted context omitted.

I think the problem is that agents are inherently stochastic. Their idea of simplification changes from message to message because whatever objective it’s operating on internally is inherently opaque and changes. No matter how much you prompt it, eventually it’s going to not do what you want it to do. I built https://github.com/thempatel/mdlr for precisely this reason: externalize the objective and force the agent to…

Interesting, I'll be testing your tool on my repos. You should publish to crates.io!

Thank you so much! I am open to any and all feedback. Please file an issue or discussion if you have things you'd like to share.

Getting this onto crates.io is a great suggestion, I will look into that!

Re: Claude Fable 5: mid-tier results on coding tasks

#242

Earlier quoted context omitted.

You don’t even know what we’re talking about in this thread, do you? We’re talking about whether corporations are going to risk using LLMs in their codebase because of the theoretical legal risk that they might produce something that would fall under open source licenses, and be difficult to untangle later. Regardless of what you think the morality is here, or what the legal situation turns out to be, this is already…

Why repeat what you already said with more words, as if I can't read, only to leave out the bit that I responded to? > we’re never going back. As a prediction, this is worthless. If everybody thinks as you do, we won't, if nobody does, we will. So yes, this is purely about morality.

It's not just about collective agreement, there's a prisoner's dilemma in there.

If some segment of engineers uses agents and outperforms engineers who don't use agents, market forces will push all other engineers to use it over time. The only way we're going back is if we get concrete evidence that engineers using agents perform worse than engineers that don't, and that evidence isn't invalidated by improved models.

Re: Claude Fable 5: mid-tier results on coding tasks

#243
post #98

Earlier quoted context omitted.

I don't understand how some of y'all use these things. I get garbage unless I give them very specific concrete tasks with as much context as possible. Anything that takes more than 30 min is usually a waste because the scope was too large.

Different people just have different concepts of what's garbage and what's not. There seems to be some kind of AI hysteria going on, with people becoming so enamoured with the AI that they accept anything it produces as if it's some gift from the gods, while others just reject it prima-facie. For example, the worst design I have seen recently was from a designer who pivoted into "vibe coding influencer". The worst co…

“One man’s trash is another man’s treasure.” takes a new meaning in today’s agentic coding world.

Re: Claude Fable 5: mid-tier results on coding tasks

#244

Earlier quoted context omitted.

> Burned $2K to see how it will perform on frontend tasks and backend tasks. When I read such statements on HN, I nearly always ask myself: if the person has such an amount of money to burn, don't there exist much more fun opportunities to burn buckets of money than doing such experiments on LLMs?

> if the person has such an amount of money to burn, don't there exist much more fun opportunities to burn buckets of money than doing such experiments on LLMs? Do you think US$2,000 is a lot of money?

That's a monthly mortgage payment for anyone who bought a starter house in a tier2/3 city prior to ~2024

Re: Claude Fable 5: mid-tier results on coding tasks

#245
post #9

This matches my experience. Burned $2K to see how it will perform on frontend tasks and backend tasks. Frontend did a significantly better job than Opus on toy-scale wireframe projects by using gimmicks like fluid dynamics. Then when given medium to big tasks like multi-page web app where layouts and aesthetics must be decided by model itself, results by Fable and Opus scored indistinguishable score from human judges…

> Burned $2K to see how it will perform on frontend tasks and backend tasks. When I read such statements on HN, I nearly always ask myself: if the person has such an amount of money to burn, don't there exist much more fun opportunities to burn buckets of money than doing such experiments on LLMs?

Yeah, seriously. I've decided not to try Fable at all, because if it is good, I don't want to get hooked, and then feel tempted to spend extra money for it when Anthropic pulls it from subscription plan access.

I'm lucky that $2k isn't a lot of money for me, though I'd much rather spend it on basically anything other than LLM credits.

As another poster noted, imagine if that money went to open source, on the regular... As an open source maintainer myself, that line of thought makes me sad.

But hey, I know I probably spend money on stuff other people would think is stupid, so I shouldn't criticize.

Re: Claude Fable 5: mid-tier results on coding tasks

#246
post #57

Earlier quoted context omitted.

Agree with this. Strange to me to frame the "training recall" as cheating (33 of the 38 cheating instances). Most people think of "cheating" as breaking rules. How is the LLM model supposed to not use what was put into the weights?

While I probably wouldn't classify it as cheating, it is an even bigger signal of concern for model quality. Cheating by breaking the rules at least implies some learned patterns. Repeating training data verbatim for narrow cases like this implies that the model is overfitting.

If we're evaluating a person, rote recall is not necessarily cheating. It's expected, but then you'd expect them to apply that rote-memorized information in a novel way later on and prove they understand how they applied their priors to the new situation.

Models don't actually reason in the same sense, so recalling rote from their training data is "cheating" in the sense that the training data cheated, not the model. So many of those benches have snaked their way into training data to make them less useful benchmarks. That, I think, is going to be a long-term difficulty in quantitatively assessing model quality and "intelligence." So it is cheating, in a sense of what we expect from the models and training data, but not in a human sense.

Re: Claude Fable 5: mid-tier results on coding tasks

#247
post #130

Earlier quoted context omitted.

I've had Fable add Chinese characters to our conversation for no reason.

Could it be that Anthropic is using the Chinese characters trick to consume less tokens behind the scenes?

It used a chinese character instead of the word "true"

Re: Claude Fable 5: mid-tier results on coding tasks

#248
post #130

Earlier quoted context omitted.

Fable is a lot like Opus at its best. It's simply more reliable and feels a bit smarter. For my use cases, using it feels very nice , and notably better than Opus. It needs less direct guidance to get reasonable looking code and I don't have to watch it as closely. For context, my Claude Code working style is quite heavy on discussion "to align" before implementing anything. We also use a good amount of Markdowns. Oh…

I've had Fable add Chinese characters to our conversation for no reason.

I've also had Fable successfully build a text editor (quill integration) into a Vaadin project that randomly loses its content after you type a few characters (this is on the 3rd iteration).

Re: Claude Fable 5: mid-tier results on coding tasks

#249

Earlier quoted context omitted.

> Burned $2K to see how it will perform on frontend tasks and backend tasks. When I read such statements on HN, I nearly always ask myself: if the person has such an amount of money to burn, don't there exist much more fun opportunities to burn buckets of money than doing such experiments on LLMs?

> if the person has such an amount of money to burn, don't there exist much more fun opportunities to burn buckets of money than doing such experiments on LLMs? Do you think US$2,000 is a lot of money?

Perhaps not for you and me (though I'm certainly not going to light $2k on fire in an LLM for shits and giggles; I have plenty of significantly better uses for that), but $2k for the vast majority of people in the US is a super big deal amount of money. Many people in the US don't even have that much to spare for an emergency, let alone for something fun.

Re: Claude Fable 5: mid-tier results on coding tasks

#250
post #122
post #115

Earlier quoted context omitted.

> that basically consisted of applying the same steps and rules n times. Why use a non-deterministic, possibly hallucinatory, definitely expensive, LLM when it sounds like a codemod is the perfect solution for this?

In this case, handling all the edge cases and variants, and testing a codemod, would have taken significantly more of my time, which costs quite a bit more than the LLM. Obviously, a deterministic tool is preferable in general, but it is not always worth bothering with for a one off task.

In that case, your original description of "basically consisted of applying the same steps and rules n times" was misleading.
Post reply on HN