Thanks for the research! Though I feel like industry veterans (especially those working with LLMs) came to this conclusion without having to write a single prompt. Even ignoring the technical merits of these kinds of hacks, if you think you've outwitted billions of dollars of statistics with a prompt, you're probably wrong at this point. What I find most interesting is the popularity of these snake oils, especially t…
I benchmarked Claude Code's caveman plugin against "be brief."
31–40 of 71 posts
Re: I benchmarked Claude Code's caveman plugin against "be brief."
#32Author here. Caveman is a popular Claude Code plugin that compresses Claude's responses via a custom skill with intensity modes. I wanted to know whether it actually beats the simplest possible alternative, prepending "be brief." to prompts. 24 prompts, 5 arms, judged by a separate Claude against per-prompt rubrics covering required facts, required terms, and dangerous wrong claims to avoid. 120 scored responses, 100…
Re: I benchmarked Claude Code's caveman plugin against "be brief."
#33Re: I benchmarked Claude Code's caveman plugin against "be brief."
#34Thanks for the research! Though I feel like industry veterans (especially those working with LLMs) came to this conclusion without having to write a single prompt. Even ignoring the technical merits of these kinds of hacks, if you think you've outwitted billions of dollars of statistics with a prompt, you're probably wrong at this point. What I find most interesting is the popularity of these snake oils, especially t…
Re: I benchmarked Claude Code's caveman plugin against "be brief."
#35Re: I benchmarked Claude Code's caveman plugin against "be brief."
#36Thanks for the research! Though I feel like industry veterans (especially those working with LLMs) came to this conclusion without having to write a single prompt. Even ignoring the technical merits of these kinds of hacks, if you think you've outwitted billions of dollars of statistics with a prompt, you're probably wrong at this point. What I find most interesting is the popularity of these snake oils, especially t…
Maybe we need a term such as prompt homeopathy to call out prompt engineering without any empirical proof.
Re: I benchmarked Claude Code's caveman plugin against "be brief."
#37Thanks for the research! Though I feel like industry veterans (especially those working with LLMs) came to this conclusion without having to write a single prompt. Even ignoring the technical merits of these kinds of hacks, if you think you've outwitted billions of dollars of statistics with a prompt, you're probably wrong at this point. What I find most interesting is the popularity of these snake oils, especially t…
Re: I benchmarked Claude Code's caveman plugin against "be brief."
#38Re: I benchmarked Claude Code's caveman plugin against "be brief."
#39Caveman made me laugh and that, in theory, should count for something.
Re: I benchmarked Claude Code's caveman plugin against "be brief."
#40Caveman is useless for me. We are in the year 2026, computers are here to serve me, and bring me comfort. Caveman is a caveman, speaks like an idiot. I don't want to interact with an idiot. It's irritating, and as the article states, an overhyped turd. It is the same idiocy that permeates EV cars. You buy an expensive car to go from A to B and at the same time offer you comfort. When I have to think about using the s…
Of the things you could complain about in modern cars as being too complicated, you chose turning on seat heating??? Like you push the seat heating button if your seat feels cold. What is there to think about?