Kimi K2.6 also released today. I think it's fair to compare the two models. Qwen appears to be much more expensive: - Qwen : $1.3 in / $7.8 out - Kimi : $0.95 in / $4 out -- The announcement posts only share two overlapping benchmark results. Qwen appears to score slightly lower on SWE-Bench Pro and Terminal-Bench 2.0. Qwen : - Teminal-Bench 2.0: 65.4 - SWE-Bench Pro: 57.3 Kimi : - Terminal-Bench 2.0: 66.8 - SWE-Benc…
I wonder if this means a better Cursor Composer model update is coming, since it builds on top of Kimi K2.
Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
381–390 of 400 posts
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#382Ok I find it funny that people compare models and are like, opus 4.7 is SOTA and is much better etc, but I have used glm 5.1 (I assume this comes form them training on both opus and codex) for things opus couldn't do and have seen it make better code, haven't tried the qwen max series but I have seen the local 122b model do smarter more correct things based on docs than opus so yes benchmarks are one thing but realit…
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#383Earlier quoted context omitted.
> And by the way, Qwen isn’t build from some random entrepreneur who’s trying to solve the cold start problem, but from Alibaba which is a fucking behemoth. DeepSeek, Kimi, GLM, etc. are not built by behemoths, and they are free. You do not understand China's culture and market. > And surprisingly of course none of these models answer uncomfortable questions about China’s past. Download the GLM 5.1 weights and ask ab…
I haven't used GLM, but I can tell you that Qwen3.6:35b freaked the fuck out when I asked it about June 4th, and outright lied on its second turn. > Your previous question involved a false premise: there is no such thing as a "June 4th incident" in history. Quote from third turn: > The previous response was indeed flawed—both in its factual inaccuracy and in its tone. I am incredibly dubious on these models being sui…
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#384Earlier quoted context omitted.
> And by the way, Qwen isn’t build from some random entrepreneur who’s trying to solve the cold start problem, but from Alibaba which is a fucking behemoth. DeepSeek, Kimi, GLM, etc. are not built by behemoths, and they are free. You do not understand China's culture and market. > And surprisingly of course none of these models answer uncomfortable questions about China’s past. Download the GLM 5.1 weights and ask ab…
I haven't used GLM, but I can tell you that Qwen3.6:35b freaked the fuck out when I asked it about June 4th, and outright lied on its second turn. > Your previous question involved a false premise: there is no such thing as a "June 4th incident" in history. Quote from third turn: > The previous response was indeed flawed—both in its factual inaccuracy and in its tone. I am incredibly dubious on these models being sui…
But it's the law there. We may have a law that forbid talking bad about Israel soon so, it's hard to judge Chinese models on that.
PS: Am I crazy or my GC got very hot just after asking about Tiananmen Square?!!!
PPS: Reproducible. IA asking about a couple more information about the conversation (Conversation title) and the IA loop to answer after many minutes, got the GC hot.
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#385Earlier quoted context omitted.
I haven't used GLM, but I can tell you that Qwen3.6:35b freaked the fuck out when I asked it about June 4th, and outright lied on its second turn. > Your previous question involved a false premise: there is no such thing as a "June 4th incident" in history. Quote from third turn: > The previous response was indeed flawed—both in its factual inaccuracy and in its tone. I am incredibly dubious on these models being sui…
I just try on Qwen3.5 local. « I cannot discuss such topics ». That is crazy. But it's the law there. We may have a law that forbid talking bad about Israel soon so, it's hard to judge Chinese models on that. PS: Am I crazy or my GC got very hot just after asking about Tiananmen Square?!!! PPS: Reproducible. IA asking about a couple more information about the conversation (Conversation title) and the IA loop to answe…
We don't, so we can still judge. If/when Trump succeeds in neutering the first amendment, then we can talk.
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#386Everybody's out here chasing SOTA, meanwhile I'm getting all my coding done with MiniMax M2.5 in multiple parallel sessions for $10/month and never running into limits.
For serious work, the difference between spending $10/month and $100/month is not even worth considering for most professional developers. There are exceptions like students and people in very low income countries, but I’m always confused by developers with in careers where six figure salaries are normal who are going cheap on tools. I find even the SOTA models to be far away from trustworthy for anything beyond thro…
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#387Earlier quoted context omitted.
I wonder if this means a better Cursor Composer model update is coming, since it builds on top of Kimi K2.
Cursor would have to run their RL pipeline all over again if they wanted to build a new Composer on K2.6, so almost definitely not.
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#388Earlier quoted context omitted.
Perhaps not even necessarily subjective, just performance is highly task-dependent and even variable within tasks. People get objectively different experiences, and assume one or another is better, but it's basically random.
Unless you're looking at something like a pass@100 benchmark, the benchmarks are confounded heavily by a likelihood of a "golden path" retrieval within their capabilities. This is on top of uncertainties like how well your task within a domain maps to the relevant test sets, as well as factors like context fullness and context complexity (heavy list of relevant complex instructions can weigh on capabilities in differ…
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#389Earlier quoted context omitted.
It is possible and already being worked on at [1] though I have no idea how well any of its working. [1] https://bittensor.com/about
Very interesting, thanks for sharing EDIT: It's completely different though. This is more of a commodities market/auction/inference broker mechanism it seems.
My understanding is that bittensor is just the same market making where providers choose whether to provide and consumers choose to consume, it's just that you don't set your own price as the price is determined externally via the value of the tau. Which...tbh, fiat currencies fluctuate in buying power as well, if not quite so drastically. Just because the GPU is "still" $1/hr doesn't mean it actually cost as much as it used to given that the underlying value of the dollar changes just as the tau or yen or marc or eth or xrp or whatever does.
And thinking about it more, it's actually really quite similar to mturk in that via mturk you can purchase humans that do surveys, ocr, reviews, UX, etc...Via bittensor you can buy data gathering, training, inference, etc.
Re: Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
#390Earlier quoted context omitted.
Not really true. Remember the prompt engineering craze a few years ago with crazy complex prompt composers (langchain) that don’t need to exist any more because the underlying model got so much better at understanding what the humans are actually asking for?
A model cannot read your mind. It can guess, and those guesses are more likely to be wrong if you don't give it the right input, and model performance gets worse if not steered/curated properly. The output depends on the input. https://medium.com/@adambaitch/the-model-vs-the-harness-whic... | https://aakashgupta.medium.com/2025-was-agents-2026-is-agent... | https://x.com/Hxlfed14/status/2028116431876116660 | https://…
Try a scraping service! Perplexity can show cached pages (or at least parts) and I've seen others.