Earlier quoted context omitted.
Deepseek is one of the worst in terms of hallucination rate according to artificial analysis' benchmark: https://artificialanalysis.ai/?omniscience=omniscience-hallu...
That's interesting. The "best" current models, Fable and especially GPT 5.6, are also lying liars that lie all the time. Seems like we're going the wrong way on hallucination.
Qwen 3.8
571–580 of 793 posts
Re: Qwen 3.8
#572I really like Qwen, even the Q2KP quants of 3.6 27B have genuinely impressive local performance on a 24GB card. It has been good enough that I am happily giving them $60 right now to try this instead of waiting to try a slightly lesser version locally. Was there ever an explanation for why we never got the weights of 3.7? I would like sourced quotes and not weird/cringe accusative speculation about distillation, or y…
do you mind sharing your settings? i just picked up an r9700 to start playing with local qwen3.6 27b and your setup sounds promising and efficient on 24gb.
As far as the model settings go I just follow what's on the card.
I'm using HauhauCS's models, they seem to do a slightly better job with their "P" quants. Especially wrt patching them to eliminate the "doom loops" that will time out the GPU (esp if you have not already given it the extra power budget, set fans to max, and lowered your max clock by about 8% to save yourself a crash/reboot).
Without getting into the weeds, unsloth covers a lot more ground so, when they're good they're great, but I've also had the most problems with them. Basically, don't be afraid to shop around and fuck with sliders.
These days it should almost always be enough to open up LMStudio, set context to max, K/V quant to 16 or 8, and off you go. I'm using a 7900XTX, with 128GB of memory for the cases where things don't fit. The default settings should be fine otherwise.
I don't know anything about the r9700, seems neat. Looks like it might play nicer with the vulkan backend than rocm, and you might have to `set GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1`. Perf should be at least as good as what I'm getting (180+ tk/s prompt, >50tk/s response), which imho is just fast enough that I can read it as it reasons and responds.
You can probably get more aggressive with KV quant (any of the 4's) and bump up the batch sizes (4096/1024 vs 2048/512, etc), it's quite situational. These days there generally should not be any serious loss of accuracy from the former. Keeping the rig cool will likely matter more, so turn your music up to hide the fans and set your rig on top of the AC exhaust lol. Oh and I guess make sure ReBAR is enabled and working, adrenaline should let you know if it's not enabled.
Re: Qwen 3.8
#573Earlier quoted context omitted.
There is clearly an anti-China bias here. Show me comments demonstrating the same level of distrust against Google for open-sourcing projects like Tensorflow, Kubernetes, Flutter, Chromium, etc.
The chromium example is wild. There's an extreme distrust and contempt for chromium becoming the defacto browser and therefore Google becoming the defacto gatekeeper of the web.
Re: Qwen 3.8
#574Earlier quoted context omitted.
> It's hard to say what their motivation is. Feels pretty easy to me. They want to turn LLMs into a commodity, and watch the US AI labs crash and burn. There will still be plenty of customers who will pay them to host the models and run inference, even if the weights are open and others can offer competing products. (If necessary, the Chinese government can ban use of foreign inference services by Chinese citizens an…
I feel like there could also be a simpler explanation. Why does a debian contributor make debian free, why do they work on this thing anyone can use? Is it because linux and debian hate windows and iOS and want to see american fail? No, it's because most debian contributors believe software source code, information, should be free, users should be free to modify the code they use, and that they're building a thing th…
Re: Qwen 3.8
#575Earlier quoted context omitted.
It is hilarious to see people from arguable the most polarized political systems in the world believing the evil 1.5 billion people across the sea share one single mind, either a saint, or a devil.
It will be a great day for China and the world when the Chinese people are free from a totalitarian dictatorship. But until then we have to speak of the policy of the Chinese government as the policy of China, even if many, or most disagree with those policies.
Re: Qwen 3.8
#576https://trilogyai.substack.com/p/qwen-38-max-benchmark-how-i...
Re: Qwen 3.8
#577Earlier quoted context omitted.
I feel like there could also be a simpler explanation. Why does a debian contributor make debian free, why do they work on this thing anyone can use? Is it because linux and debian hate windows and iOS and want to see american fail? No, it's because most debian contributors believe software source code, information, should be free, users should be free to modify the code they use, and that they're building a thing th…
Models are treated as weapons with export controls - if they can do this it’s with the blessing of the Chinese government who’s getting something out of it. It’s fairly obviously about being a nuisance to the US.
Re: Qwen 3.8
#578Earlier quoted context omitted.
That's interesting. The "best" current models, Fable and especially GPT 5.6, are also lying liars that lie all the time. Seems like we're going the wrong way on hallucination.
And I remember Sam Altman saying 2 years ago or more, that hallucination was already "fixed" internally. Well, clearly not.
I am surprised by how much the really giant models hallucinate, though. My vague feeling was that little models hallucinate a lot because they just don't know anything (the world's knowledge simply does not fit in a few GB) and don't know how to say, "I don't know". But, the big models kinda do know everything, and yet, here we are, they're still making shit up all the time.
Re: Qwen 3.8
#579> "compatible" instead of "comparable." It boggles my mind how you can train a frontier model but not write a tweet without an obvious typo.
Re: Qwen 3.8
#580Earlier quoted context omitted.
I feel like there could also be a simpler explanation. Why does a debian contributor make debian free, why do they work on this thing anyone can use? Is it because linux and debian hate windows and iOS and want to see american fail? No, it's because most debian contributors believe software source code, information, should be free, users should be free to modify the code they use, and that they're building a thing th…
Yeah the Chinese totally have a really good history with being completely open and giving lol. The Chinese government totally has not been hacking into American and Western fortune 500 companies for the past few decades stealing R&D and tech to use for themselves. The Chinese also totally do not steal hundreds of billions of dollars of IP from America annually. Totally not something they would do!