Live data from Hacker News

Without benchmarking LLMs, you're likely overpaying

karllorey.com

71–80 of 100 posts

Re: Without benchmarking LLMs, you're likely overpaying

#71
I'm consistently amazed at how much some individuals spend on LLMs.

I get a good amount of non-agentic use out of them, and pay literally less than $1/month for GLM-4.7 on deepinfra.

I can imagine my costs might rise to $20-ish/month if I used that model for agentic tasks... still a very far cry from the $1000-$1500 some spend.

Re: Without benchmarking LLMs, you're likely overpaying

#72
post #66

Earlier quoted context omitted.

I hate thumbs up/down. 2 values is too little. I understand that 5 was maybe too much, but thumbs up/down systems need an explicit third "eh, it's okay" value for things I don't hate, don't want to save to my library, but I would like the system to know I have an opinion on. I know that consuming something and not thumbing it up/down sort-of does that, but it's a vague enough signal (that could also mean "not close e…

Here's the discussion from back in the day when this changed: https://news.ycombinator.com/item?id=837698 In practice, people generally didn't even vote with two options, they voted with one! IIRC youtube did even get rid of downvotes for a while, as they were mostly used for brigading.

> IIRC youtube did even get rid of downvotes for a while, as they were mostly used for brigading.

No, they got rid of them most likely because advertisers complained that when they dropped some flop they got negative press from media going "lmao 90% dislike rate on new trailer of ".

Stuff disliked to oblivion was either just straight out bad, wrong (in case of just bad tutorials/info) and brigading was very tiny percentage of it.

Re: Without benchmarking LLMs, you're likely overpaying

#73
post #3

I'd second this wholeheartedly Since building a custom agent setup to replace copilot, adopting/adjusting Claude Code prompts, and giving it basic tools, gemini-3-flash is my go-to model unless I know it's a big and involved task. The model is really good at 1/10 the cost of pro, super fast by comparison, and some basic a/b testing shows little to no difference in output on the majority of tasks I used Cut all my sub…

LLM bubble will burst the second investors figure out how much well managed local model can do

Re: Without benchmarking LLMs, you're likely overpaying

#74
post #71

I'm consistently amazed at how much some individuals spend on LLMs. I get a good amount of non-agentic use out of them, and pay literally less than $1/month for GLM-4.7 on deepinfra. I can imagine my costs might rise to $20-ish/month if I used that model for agentic tasks... still a very far cry from the $1000-$1500 some spend.

Doesn't this depend a lot on private vs company usage? There's no way I could spend more than a few hundreds alone, but when you run prompts on 1M entities in some corporate use case, this will incur costs, no matter how cheap the model usage.

Re: Without benchmarking LLMs, you're likely overpaying

#75
post #59

I paid a total of 13 US Dollars for all my llm usage in about 3 years. Should I analyze my providers and see if there's room for improvement?

How? All LLM-as-a-Servive's are prohibitively expensive for me. $13 over 3 years sounds too-good-to-be-true.

All local CLIs with free to use models. CLIs are opencode, iflow, qwen, gemini.

What I did splurge on was brief openai access for some subtitle translator program and when I used the deepseek api. Actually I think that $13 includes some as yet unused credits. :D

I'd be happy to provide details if CLIs are an option and you don't m ind some sweatshop agent. :)

(I am just now noticing I meant to type 2 years not 3 above. Sorry about that.)

Re: Without benchmarking LLMs, you're likely overpaying

#77

Earlier quoted context omitted.

Here's the discussion from back in the day when this changed: https://news.ycombinator.com/item?id=837698 In practice, people generally didn't even vote with two options, they voted with one! IIRC youtube did even get rid of downvotes for a while, as they were mostly used for brigading.

> IIRC youtube did even get rid of downvotes for a while, as they were mostly used for brigading. No, they got rid of them most likely because advertisers complained that when they dropped some flop they got negative press from media going "lmao 90% dislike rate on new trailer of ". Stuff disliked to oblivion was either just straight out bad, wrong (in case of just bad tutorials/info) and brigading was very tiny perc…

YouTube never got rid of downvotes they just hid the count. Channel admins can still see it and it still affects the algorithm

Re: Without benchmarking LLMs, you're likely overpaying

#78

Earlier quoted context omitted.

Here's the discussion from back in the day when this changed: https://news.ycombinator.com/item?id=837698 In practice, people generally didn't even vote with two options, they voted with one! IIRC youtube did even get rid of downvotes for a while, as they were mostly used for brigading.

> IIRC youtube did even get rid of downvotes for a while, as they were mostly used for brigading. No, they got rid of them most likely because advertisers complained that when they dropped some flop they got negative press from media going "lmao 90% dislike rate on new trailer of ". Stuff disliked to oblivion was either just straight out bad, wrong (in case of just bad tutorials/info) and brigading was very tiny perc…

Oh, didn't they remove the dislike count after people absolutely annihilated one of their yearly rewind with dislikes?

Re: Without benchmarking LLMs, you're likely overpaying

#79
post #66

Earlier quoted context omitted.

I hate thumbs up/down. 2 values is too little. I understand that 5 was maybe too much, but thumbs up/down systems need an explicit third "eh, it's okay" value for things I don't hate, don't want to save to my library, but I would like the system to know I have an opinion on. I know that consuming something and not thumbing it up/down sort-of does that, but it's a vague enough signal (that could also mean "not close e…

Here's the discussion from back in the day when this changed: https://news.ycombinator.com/item?id=837698 In practice, people generally didn't even vote with two options, they voted with one! IIRC youtube did even get rid of downvotes for a while, as they were mostly used for brigading.

Youtube always kept downvotes and the 'dislike' button, the change (which still applies today) was that they stopped displaying the downvote count to users - the button never went away though.

Visit a youtube video today, you can still upvote and downvote with the exact same thumbs up or down, the site however only displays to you the count of upvotes. The channel owners/admins can still see the downvote count and the downvotes presumably still inform YouTube's algorithms.

Re: Without benchmarking LLMs, you're likely overpaying

#80

Earlier quoted context omitted.

You use stuff from xAi and Elmo? I'm unwilling to look past Musk's politics, immorality, and manipulation on a global scale

Grok is the best general purpose LLM in my experience. Only Gemini is comparable. It would be silly to ignore it, and xAI is less evil than Google these days.

When's the last time Sundar Pichai did a Hitler salute or had his creation calling itself "Mecha Hitler"?
Post reply on HN