More benchmaxxing I see. Too bad there’s no rig with 256gb unified ram for under $1000
do you know if they did this to it? https://research.google/blog/turboquant-redefining-ai-effici...
So a quantized KV cache now must see less degradation
41–50 of 563 posts
More benchmaxxing I see. Too bad there’s no rig with 256gb unified ram for under $1000
do you know if they did this to it? https://research.google/blog/turboquant-redefining-ai-effici...
So a quantized KV cache now must see less degradation
Earlier quoted context omitted.
It's a MoE model and the A3B stands for 3 Billion active parameters, like the recent Gemma 4. You can try to offload the experts on CPU with llama.cpp (--cpu-moe) and that should give you quite the extra context space, at a lower token generation speed.
Mac has unified memory, so 36GB is 36GB for everything- gpu,cpu.
Does anyone have any experience with Qwen or any non-Western LLMs? It's hard to get a feel out there with all the doomerists and grifters shouting. Only thing I need is reasonable promise that my data won't be used for training or at least some of it won't. Being able to export conversations in bulk would be helpful.
The Chinese models are generally pretty good. > Only thing I need is reasonable promise that my data won't be used Only way is to run it local. I personally don’t worry about this too much. Things like medical questions I tend to do against local models though
"open source" give me the training data?
"open source" give me the training data?
#include
int m
I get nonsensical autocompletions like: #include
int m
What is going on?I don't want "Agentic Power". I want to reduce AI to zero. Granted, this is an impossible to win fight, but I feel like Don Quichotte here. Rather than windmill-dragons, it is some skynet 6.0 blob.