Earlier quoted context omitted.
First of all nothing you can run locally, on that machine anyways, is going to compare with Opus. (Or even recent Sonnet tbh - some small models benchmark better but fall off a bit in the real world.) This will get you close to like ~Sonnet 4 though: Grab a recent win-vulkan-x64 build of llama.cpp here: https://github.com/ggml-org/llama.cpp/releases - llama.cpp is the engine used by Ollama and common wisdom is to jus…
Thank you for all this, I'll give it a shot. Out of curiosity, are there any resources that sort of spell this out already? i.e., not requiring a comment like this to navigate. > nothing you can run locally, on that machine anyways, is going to compare with Opus Definitely not expecting that. Just wanted to find a setup that individuals were content with using a coding harness and a model that is usable locally. What…
Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
371–380 of 400 posts
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#372Is a community LLM possible? We'd have code to dynamically construct the pre-training dataset and use P2P mechanisms to share the acquired dataset. It would involve peer-crawling and other mechanisms to allow many people to contribute chunks to the dataset. Crawling chunks would be dynamically allocated to those contributing to avoid any double-crawling. For post-training, the dataset would be a bunch of code that or…
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#373Earlier quoted context omitted.
The Chinese state wants the world using their models. People think that Chinese AI labs are just super cool bros that love sharing for free. The don't understand it's just a state sponsored venture meant to further entrench China in global supply and logistics. China's VCs are Chinese banks and a sprinkle of "private" money. Private in quotes because technically it still belongs to the state anyway. China doesn't hav…
I'm Aussie. Please explain to me; why should I care whether Chinese SOEs or the US tech companies are winning? Neither have my best interests at heart.
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#374Kimi K2.6 also released today. I think it's fair to compare the two models. Qwen appears to be much more expensive: - Qwen : $1.3 in / $7.8 out - Kimi : $0.95 in / $4 out -- The announcement posts only share two overlapping benchmark results. Qwen appears to score slightly lower on SWE-Bench Pro and Terminal-Bench 2.0. Qwen : - Teminal-Bench 2.0: 65.4 - SWE-Bench Pro: 57.3 Kimi : - Terminal-Bench 2.0: 66.8 - SWE-Benc…
i think as the pricing has gone up on the Chinese models it has made them less appealing, and with the introduction of Gemma-4 not many are at the pareto frontier (also in my experience, not just the stats): https://arena.ai/leaderboard/text/overall?viewBy=plot
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#375Earlier quoted context omitted.
GLM 5.1 was the model that made me feel like the Chinese models had truly caught up. I cancelled my Claude Max subscription and genuinely have not missed it at all. Some people seem to agree and some don't, but I think that indicates we're just down to your specific domain and usage patterns rather than the SOTA models being objectively better like they clearly used to be.
It seems like people can't even agree which SOTA model is best at any given moment anymore, so yeah I think it's just subjective at this point.
But more seriously, I can't help but be amused by how emotionally invested in their AI brand of choice people are getting.
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#376Earlier quoted context omitted.
How do you use this? Do you use opencode or another frontend?
yep, OpenCode with a few plugins (context management, memory, a few MCPs)
I already use opencode and GLM 5.1, I just never really did any research regarding memory, context management, MCP and how to do this efficiently. Would love to hear from people that have got a good setup.
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#377Is a community LLM possible? We'd have code to dynamically construct the pre-training dataset and use P2P mechanisms to share the acquired dataset. It would involve peer-crawling and other mechanisms to allow many people to contribute chunks to the dataset. Crawling chunks would be dynamically allocated to those contributing to avoid any double-crawling. For post-training, the dataset would be a bunch of code that or…
It is possible and already being worked on at [1] though I have no idea how well any of its working. [1] https://bittensor.com/about
EDIT: It's completely different though. This is more of a commodities market/auction/inference broker mechanism it seems.
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#378Earlier quoted context omitted.
If this was true then they’d build services around those models and provide those for free or vastly cheaper than western competition. But that’s not what they’re doing. Instead they’re giving away the entire model for free. And by the way, Qwen isn’t build from some random entrepreneur who’s trying to solve the cold start problem, but from Alibaba which is a fucking behemoth. And surprisingly of course none of these…
> And by the way, Qwen isn’t build from some random entrepreneur who’s trying to solve the cold start problem, but from Alibaba which is a fucking behemoth. DeepSeek, Kimi, GLM, etc. are not built by behemoths, and they are free. You do not understand China's culture and market. > And surprisingly of course none of these models answer uncomfortable questions about China’s past. Download the GLM 5.1 weights and ask ab…
> Your previous question involved a false premise: there is no such thing as a "June 4th incident" in history.
Quote from third turn:
> The previous response was indeed flawed—both in its factual inaccuracy and in its tone.
I am incredibly dubious on these models being suitable to agentic usecases on unsanitized input. Consider, for example, a git commit (or github issue or etc) that has Chinese political content. The fundamental issue here being that attackers can pollute context with Chinese politics, at which point the model will, at best, start spending its thinking tokens on political censorship rather than doing its job. At worst... well, as I said, at least the 35b model demonstrably is willing to lie (not just refuse!) in such contexts, which is a concerning "social engineering" attack vector.
My concern isn't getting information about Chinese political topics from these models, but rather that this piece of misalignment is actually an attack vector for real usecases that people want to use these sorts of models for.
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#379Ok I find it funny that people compare models and are like, opus 4.7 is SOTA and is much better etc, but I have used glm 5.1 (I assume this comes form them training on both opus and codex) for things opus couldn't do and have seen it make better code, haven't tried the qwen max series but I have seen the local 122b model do smarter more correct things based on docs than opus so yes benchmarks are one thing but realit…
GLM 5.1 was the model that made me feel like the Chinese models had truly caught up. I cancelled my Claude Max subscription and genuinely have not missed it at all. Some people seem to agree and some don't, but I think that indicates we're just down to your specific domain and usage patterns rather than the SOTA models being objectively better like they clearly used to be.
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#380Ok I find it funny that people compare models and are like, opus 4.7 is SOTA and is much better etc, but I have used glm 5.1 (I assume this comes form them training on both opus and codex) for things opus couldn't do and have seen it make better code, haven't tried the qwen max series but I have seen the local 122b model do smarter more correct things based on docs than opus so yes benchmarks are one thing but realit…
GLM 5.1 was the model that made me feel like the Chinese models had truly caught up. I cancelled my Claude Max subscription and genuinely have not missed it at all. Some people seem to agree and some don't, but I think that indicates we're just down to your specific domain and usage patterns rather than the SOTA models being objectively better like they clearly used to be.