Live data from Hacker News

Qwen 3.8 27B

huggingface.co

441–450 of 848 posts

Re: Qwen 3.8 27B

#441

Earlier quoted context omitted.

[flagged]

Yeah I'm sure everyone on r/cursor or in previous HN threads about Grok 4.5 or 4.6 are all unserious and insane. No one actually cares about the politics as long as the model codes well. Edit, quite interesting to see the reception to this comment compared to essentially the same type of comment I made on a Grok 4.6 benchmark HN post: https://news.ycombinator.com/item?id=49275385#49275571 It's true that Cursor gives…

"No one actually cares about the politics as long as the model codes well."

> I totally care about politics, especially when it comes to not giving my money to people like musk.

Re: Qwen 3.8 27B

#442

Earlier quoted context omitted.

I'm still confused about Qwen 3.6 35B A3B. Everything I read said that the 27B model performs better at coding tasks, so what's the purpose of the 35B model?

It's better for VRAM poor people. I get 4-5 t/s with 27B and 20-30 t/s with 35B A3B.

Also radically better on an M1 Max. I get well up into the 60s t/s with the A3B, stuck at 9.5 or so with this new 27B, though perhaps an MLX build will help.

Re: Qwen 3.8 27B

#443

Really excited to see what people do with this. 3.7 27B was probably the best compromise between size and intelligence to run on consumer hardware

Do you mean 3.6 27b? Because qwen 3.7 didn't have an open weight version

Re: Qwen 3.8 27B

#444

Earlier quoted context omitted.

The director receives accreditation for directing the film, not creating it.

Doesn't the director generally receive more credit than the producer? How many films do you remember the producer above the director?

What does it matter who receives more credit? The producer produces the film and the director directs it. If I pick my phone and record a video, I'm now the producer and director of the film and the sole creator of it.

If there were multiple people involved in the creation of a film I helped to create, I cannot factually say I created it. Just like if someone builds something using code generated by AI, they can't factually say they created it.

Re: Qwen 3.8 27B

#445
Beating or comparable to Opus 4.6 in benchmarks. Opus 4.6 was released in February, 2026. So if we still want to talk about a "6 month difference" between Chinese and American AI, the sentence should now be:

Chinese (small model) AI is 6 months behind American (largest model) AI.

Re: Qwen 3.8 27B

#446
post #419

Absolutely the best pelican I've seen from a model that runs on my laptop: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Bicycle is the right shape. Pelican beak is excellent. Nice background. Most importantly, the pelican has one leg on each side of the bicycle - that's very rare. (No chain on this bicycle though - in the reasoning trace it says "already chainstay... skip chain detail; maybe a smal…

What's the result if you ask it for a "pelican equipment case" ? I've been trying the anti-bicycle pelican on some LLMs and the results are much more varied than the bicycle prompt. Some have very different ideas of pelicans (you'll get a small case with one DSLR camera in it, or a long rifle case, etc). You'll also get cases that are isometric view, or flat plane view from the front, or open or closed.

Re: Qwen 3.8 27B

#447
post #109

Since it might be helpful to some, here's my current commandline for llama.cpp running on an RTX 4090 with my monitor moved to the iGPU to free up all of its VRAM. llama-server -m Qwen3.8-27B-IQ4_NL.gguf --mmproj mmproj-BF16.gguf -c 170000 --parallel 1 -ngl -1 --cache-type-k q8_0 --cache-type-v q8_0 -b 1024 -ub 512 --flash-attn on --no-context-shift --no-mmproj-offload --spec-type draft-mtp --spec-draft-n-max 5 --spe…

Jeez, llama.c++ is becoming the ffmpeg cargo cult CLI now

For the things it does, what tool is better than ffmpeg? Really struggling to see the cargo cult angle. Similar for llama.cpp, as it is literally the only framework I can get to run on my multi-Radeon rig. It is the most portable runtime out there.

Re: Qwen 3.8 27B

#448
post #68

Earlier quoted context omitted.

> I've settled on GLM-5.3 (formerly Deepseek v4 pro 0813) for architecting Dude, GLM-5.3 released _today_. The phrasing "I've settled on" is incorrect for this context.

I'm generally extremely skeptical about a lot of the model hype that show up in comments. Except when there is an extreme mismatch the performance, quirks and quality of these things are difficult to nail down. You wouldn't know that from the comment section of every single release . I think some of these are excited, eager users always ready to hype up the new thing. The same crowd that previously would constantly p…

> I think most of it is bot driven spam by the various labs. It's hard not to notice 3 month old accounts with very strong opinions about various frontier labs and little else.

I've assumed the same as well.

I also assume that many of the companies developing these models engage in benchmaxxing.

At my company we've developed our own internal benchmarks for evaluating LLM models as they become available. The benchmarks are tailored to our particular use cases but the utility and knowledge our benchmarks assess is still fairly universally applicable. I see wide differences between what our internal benchmarks report and what the major benchmarks do.

There was a whole lot of fanfare about how amazing GLM 5.2 was when it was released, but it was pure rubbish on our internal benchmark -- far behind OpenAI, Anthropic, Gemini, DeepSeek, etc. I don't know how to reconcile the fact that GLM 5.2 performed very well on some of the major public benchmarks, but consistently performs so poorly on ours. OpenAI models tend to dominate our internal benchmarks.

Re: Qwen 3.8 27B

#449

I'm struggling to figure out what to use this for. From the intelligence benchmarks in OMLX. If only they would release another MoE model. Intelligence Benchmark Comparison --- Detail --- Model: scottlowry--Qwen3.8-27B-oQ4e-mtp Benchmark Accuracy Correct Total Time(s) Think -------------------------------------------------------------- GSM8K 93.3% 28 30 282 No MATHQA 46.7% 14 30 26.3 No HUMANEVAL 96.7% 29 30 156.5 No…

Try manually asking both more discrete esoteric knowledge questions. Or use benchmarks which are less coding focused. The 3.6-35B-A3B with post-training may do well in coding type benchmarks and math but the density of its knowledge falls off in my experience (vs 3.6 27B dense Q8-K-XL unsloth GGUF) when you need to use it for less commonly used domains of knowledge.

Re: Qwen 3.8 27B

#450

Earlier quoted context omitted.

> ...but no. They do not beat opus on real-world usage. I agree, but then we just need meaningful benchmarks that clearly show that! Otherwise it's hand waving about something that should be put on paper in quantifiable terms.

A wise man once said, "not everything that counts can be counted, and not everything that can be counted counts".

That’s well and good, but how are we supposed to evaluate the accuracy of random HN comments without anything resembling somewhat objective metrics? People say all manner of things, and usually it’s contradictory. What heuristic do you propose?
Post reply on HN