The author of this post should benchmark his own blog for accessibility metrics, text contrast is dreadful.. On the other hand, this would be interesting for measuring agents in coding tasks, but there's quite a lot of context to provide here, both input and output would be massive.
Appreciate the feedback, will work on that.
Without benchmarking LLMs, you're likely overpaying
51–60 of 100 posts
Re: Without benchmarking LLMs, you're likely overpaying
#52Anecdotal tip on LLM-as-judge scoring - Skip the 1-10 scale, use boolean criteria instead, then weight manually e.g. - Did it cite the 30-day return policy? Y/N - Tone professional and empathetic? Y/N - Offered clear next steps? Y/N Then: 0.5 * accuracy + 0.3 * tone + 0.2 * next_steps Why: Reduces volatility of responses while still maintaining creativeness (temperature) needed for good intuition
Isn’t this just rubrics?
Re: Without benchmarking LLMs, you're likely overpaying
#53Re: Without benchmarking LLMs, you're likely overpaying
#54Sorry, this just makes no sense to start off with. What do you mean?
Re: Without benchmarking LLMs, you're likely overpaying
#55> it's the default: You have the API already Sorry, this just makes no sense to start off with. What do you mean?
Re: Without benchmarking LLMs, you're likely overpaying
#56Presumably that'll be some sort of funnel for a paid upload of prompts.
Re: Without benchmarking LLMs, you're likely overpaying
#57Anecdotal tip on LLM-as-judge scoring - Skip the 1-10 scale, use boolean criteria instead, then weight manually e.g. - Did it cite the 30-day return policy? Y/N - Tone professional and empathetic? Y/N - Offered clear next steps? Y/N Then: 0.5 * accuracy + 0.3 * tone + 0.2 * next_steps Why: Reduces volatility of responses while still maintaining creativeness (temperature) needed for good intuition
Re: Without benchmarking LLMs, you're likely overpaying
#58Aren't you supposed to customize the prompts to the specific models?
Re: Without benchmarking LLMs, you're likely overpaying
#59Re: Without benchmarking LLMs, you're likely overpaying
#60I love the user experience for your product. You're giving a free demo with results within 5 minutes and then encourage the customer to "sign in" for more than 10 prompts. Presumably that'll be some sort of funnel for a paid upload of prompts.
Here's a bug report, by switching the model group the api hangs in private mode.