Earlier quoted context omitted.
That kind of result makes me suspicious of benchmaxxing. Qwen 27B is 100x smaller than Opus 4.7. Is it really 100x more parameter-efficient? Two orders of magnitude is hard to believe. I don't have the hardware to run a 27B, but I'm curious what real world use is like. Maybe I'll have to buy some usage on a cloud provider to run my own tests, but this seems fishy to me.
It's very agentic coding focused; and I'd say a good executor but certainly not Opus in scale; overall knowledge; long-horizon work and recovery; etc. e.g. If you try to chat to it about something philosophical for example, or maybe a debate / creative writing, then you'll very quickly see how it is still a much smaller model at the end of the day. Still, it's such a relatively accessible model to run, and I find a b…
Qwen 3.8 27B
591–600 of 848 posts
Re: Qwen 3.8 27B
#592One thing a lot of people don't seem to factor when hyping Qwen is how much models like this tend to 'overthink' with seemingly endless 'second guessing'. 3.8 seems no different from what I've tried thus far. As capable as it is, it's hard to justify using it when a competing model (e.g. Gemma4:26b-a3b) can consistently achieve the same or similar response with only 1/10th as many 'thinking' tokens, achieve much high…
Use 3.6 27b as a daily driver for months with charmbracelet crush. Gemma 26b-A3b is not even remotely comparable in terms of coding for me. YMMV depending on how you work, what harness you use, etc I suppose.
I didn't realise there are people out there unironically using crush
Re: Qwen 3.8 27B
#593Earlier quoted context omitted.
Grok 4.6 is a game-changer. I have yet to go back to other models after starting to use it. You just can't beat the price + output quality (even K3 is more expensive)
Supporting a far-right megalomaniac, whilst helping them to train their ML, and giving them all your data ... what could go wrong.
Re: Qwen 3.8 27B
#594Earlier quoted context omitted.
Asking it factual information[1]. You just can't compress the entire human knowledge into a 30GB file. [1] Without searching the internet. And even if you allow it, you'll get much worse results because search means browsing and parsing the top results, and search results are horrible, whereas internal knowledge from training encompasses the entire internet plus all books including very niche stuff.
> You just can't compress the entire human knowledge into a 30GB file. ...I'm just asking questions here... how sure are we of this? If you'd asked me six years ago whether we could compress all of human knowledge into a 1 TB file, I would have said no, and yet here we are. If you can do 1 TB, why not 30 GB? It's, like, within one order of magnitude.
On a more serious note, it depends on your cutoff for "entire human knowledge". It's easy to prove for a generous interpretations of "entire human knowledge" that it can't be done, but hard for something like "all useful human knowledge".
Re: Qwen 3.8 27B
#595Re: Qwen 3.8 27B
#596Wow. This model is so good, and we have GLM 5.3 (seems great voor security related work) and Deepseek. In a few months we'll have Fable/Sol-like capabilities that are not coming from the big US companies. I feel as a programmer that that is more than enough. How wil OpenAI and Anthropic survive when frontier model intelligence becomes commoditized?
If you look at the hiring marketplace, being just marginally better than your peers can be very lucrative.
If you’re competing on speed or capability as a company (or as an employee), you’re probably going to be willing to pay for the frontier.
Re: Qwen 3.8 27B
#597Earlier quoted context omitted.
What are you working on? That can dictate which models are best.
Anything from ML pipelines for language specific pruning over a Rust/JS/CSS mix codebase to assistance in motorcycle maintenance and different canvas coatings. Most of my evals build on those requirements and especially past failures, whether in pure information, task execution and coding or tool calling beyond the overfitted mainstream. All stuff derived from actual failures encountered, some still only few models c…
Likewise I have a tool-use eval set and a browser-use eval set. I use frontier models for most interactive tasks that are not home assistant or background agents.
I can imagine someone could build evals for that but I have never done so.
Re: Qwen 3.8 27B
#598Re: Qwen 3.8 27B
#599Credit where it's due. Qwen 3.8 27B is only the second local model after Gemma 4 that managed to correctly reason through one of my private benchmarks. It took 5x as many tokens to do it and 12m30s with MTP enabled, but it did do it. Gemma 4 reasoned through it more implicitly, while Qwen 3.8 reasoned more explicitly. Laguna and Muse Glimmer failed hard on it, though they're useful for other tasks. The VRAM usage see…
I was quite impressed by Muse Glimmer, and while I am sure people will observe that it is less good on benchmarks, my first experiences with this new 27B have been somewhat exasperating, whereas testing Muse Glimmer was rather fun. I have not tested either in an agentic context, mind you.
Re: Qwen 3.8 27B
#600Earlier quoted context omitted.
So add a pre-ingestion agent to your Pi Coding Agent subagent repertoire, problem solved.
Woah woah woah buddy... suggesting that people use anything other than ClaudeCode or Codex is simply not allowed around these parts.