I Benchmarked Local LLMs on the Laptop I Have
1–8 of 8 posts
Re: I Benchmarked Local LLMs on the Laptop I Have
#2It just feels, at the moment, there's no task you'd want to throw at this over a frontier model, and while prices are subsidized you really can't compete at home
Re: I Benchmarked Local LLMs on the Laptop I Have
#3I ran Kimi a few days ago at 1/2token per second. I only get to play with it during the weekend when i have time, but I'm certain I'll be able to get it up to 5tk/sec when I'm done in a month or two. So yeah, we can run them all locally. Folks might say it's not run if it's that slow, but feh! If you can run the best AI model locally at 1tk/sec, why won't you?
Re: I Benchmarked Local LLMs on the Laptop I Have
#4Re: I Benchmarked Local LLMs on the Laptop I Have
#5From the article, "One guide this summer was literally titled “Open Weights You Can’t Run.”" I ran Kimi a few days ago at 1/2token per second. I only get to play with it during the weekend when i have time, but I'm certain I'll be able to get it up to 5tk/sec when I'm done in a month or two. So yeah, we can run them all locally. Folks might say it's not run if it's that slow, but feh! If you can run the best AI model…
Maybe the question is not can you, but rather should you?
Re: I Benchmarked Local LLMs on the Laptop I Have
#6This kinda goes along with my ancedotal findings "Local models are cool but still not quite there for consumer-grade hardware, but getting closer and closer" It just feels, at the moment, there's no task you'd want to throw at this over a frontier model, and while prices are subsidized you really can't compete at home
Re: I Benchmarked Local LLMs on the Laptop I Have
#7yeah you realize you can turn reasoning off on the newer the models right? also try my version of Gemma-12b https://huggingface.co/macwhisperer/Gemma4-12B-SuperDense or try my qwen 3.5-9b
Re: I Benchmarked Local LLMs on the Laptop I Have
#8From the article, "One guide this summer was literally titled “Open Weights You Can’t Run.”" I ran Kimi a few days ago at 1/2token per second. I only get to play with it during the weekend when i have time, but I'm certain I'll be able to get it up to 5tk/sec when I'm done in a month or two. So yeah, we can run them all locally. Folks might say it's not run if it's that slow, but feh! If you can run the best AI model…
Well considering I can just use DeepSeek for pennies, not sure what I might get out of that 1 tk/sec locally. Maybe the question is not can you, but rather should you?