Live data from Hacker News

Qwen 3.8 27B

huggingface.co

741–750 of 848 posts

Re: Qwen 3.8 27B

#741

Earlier quoted context omitted.

> General knowledge: I usually ask 2 questions many small models get wrong: summarize Operation Trojan Horse by John Keel, give publication year. Summarize the Ariel school incident of 1994. Qwen3.8 got the first question right, along correct publication year, gave glorious detail on Keel's theory, but got the second wrong. It thought that school was located in the USA. Ah well. I'm not sure that this means anything.…

Did the model refuse to answer? Did it say that it doesn't know? If not, then it's a fair game in my opinion.

[dead]

Re: Qwen 3.8 27B

#742

Tested the model briefly with my usual eval: a couple of questions on general knowledge most small models often get wrong, then write a fully-featured todo list app in JS, then rewrite the same app in Rust with Tauri. Granted, most models are well trained on basic todo apps, but it gives me an idea of the basic SWE capabilities I can build on. As far as I'm concerned, if it can successfully setup a local git repo, wr…

Meaningless yet fun fact: DeepSeek V4 Pro 0813 made a much worse icon for the same app, and only produced an SVG I had to convert manually to .png. Qwen3.8 made a perfect icon in .png. I don't yet know how it did it, but it did it. Qwen3.8 also seems to know French quite a bit better than Copilot, at least on common expressions. I have yet to run more tests for languages, but I'm baffled by its finer accuracy on the…

copilot isn't a model. it's a model hoster.

Re: Qwen 3.8 27B

#743
post #662

Earlier quoted context omitted.

Diverging from the sampler used in RL training is not good for long multi-turn results-- it's a great way to knock models into reasoning loops that wouldn't otherwise.

Peer reviews NeurIPS caliber paper to prove that? Because I can show you one titled "Long context generation is a sampling problem"...

Show us, we're curious. Did you upload to ArXiv yet?

Re: Qwen 3.8 27B

#744

Tested the model briefly with my usual eval: a couple of questions on general knowledge most small models often get wrong, then write a fully-featured todo list app in JS, then rewrite the same app in Rust with Tauri. Granted, most models are well trained on basic todo apps, but it gives me an idea of the basic SWE capabilities I can build on. As far as I'm concerned, if it can successfully setup a local git repo, wr…

> write a fully-featured todo list app in JS

Would you mind sharing how you prompt this? I'm not a developer myself (just someone who occasionally dabbles, though most of my coding was pre-LLMs) and curious to see how much info/instruction you consider necessary to test them making an actual app (albeit a simple one).

Re: Qwen 3.8 27B

#745
post #421
post #419

Absolutely the best pelican I've seen from a model that runs on my laptop: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Bicycle is the right shape. Pelican beak is excellent. Nice background. Most importantly, the pelican has one leg on each side of the bicycle - that's very rare. (No chain on this bicycle though - in the reasoning trace it says "already chainstay... skip chain detail; maybe a smal…

For anyone who followed yesterday's Gemini 3.7 Flash pelican which rendered in Safari but not in Firefox or Chrome... https://news.ycombinator.com/item?id=49289112#49290012 . - that turned out to be my fault, not the model. My SVG rendering software was stripping some attributes. Here's the Gemini 3.7 Flash pelicans in the fixed renderer: https://tools.simonwillison.net/markdown-svg-renderer.html#u...

>> Most importantly, the pelican has one leg on each side of the bicycle - that's very rare.

How can you say "each side"? I don't see any z-ordering between bicycle frame and legs.

Re: Qwen 3.8 27B

#746

Tested the model briefly with my usual eval: a couple of questions on general knowledge most small models often get wrong, then write a fully-featured todo list app in JS, then rewrite the same app in Rust with Tauri. Granted, most models are well trained on basic todo apps, but it gives me an idea of the basic SWE capabilities I can build on. As far as I'm concerned, if it can successfully setup a local git repo, wr…

> on an admittedly overpowered laptop: 15 tokens/s in power save mode on this MacBook M5 Max 48GB, and 30 in performance mode.

Okay, so I'm here to brag a little. I love that I also get 30 t/s on $1500 of decade-old hardware: dell r720 w 2x tesla v100s!

Re: Qwen 3.8 27B

#747

Earlier quoted context omitted.

Doesn't the director generally receive more credit than the producer? How many films do you remember the producer above the director?

What does it matter who receives more credit? The producer produces the film and the director directs it. If I pick my phone and record a video, I'm now the producer and director of the film and the sole creator of it. If there were multiple people involved in the creation of a film I helped to create, I cannot factually say I created it. Just like if someone builds something using code generated by AI, they can't fa…

There's a bit more nuance to it though, I think, because traditionally there has been more of a firm line between humans and tools.

For example, most people would agree with this line you wrote: "If I pick my phone and record a video, I'm now the producer and director of the film and the sole creator of it." A pedant could say "woah hang on, that's ignoring the fact that actually the iPhone is the one recording the images that make up the film, how can the human get all the credit", yet nobody would actually make that argument when discussing who created the film.

Generative AI is pretty much the first time (maybe there are some niche contradictions to this claim?) we consider a tool to be contributing enough creativity to the process that we don't all agree "only one person was operating this tool, so that person is the sole creator" - but some people DO still hold that line, and do consider the human who wrote the prompt to be the creator.

And I don't think there's any objective technical metric we can use to say who's right, it comes down to our collective judgement deciding where the line is.

Re: Qwen 3.8 27B

#748
post #745
post #421

Earlier quoted context omitted.

For anyone who followed yesterday's Gemini 3.7 Flash pelican which rendered in Safari but not in Firefox or Chrome... https://news.ycombinator.com/item?id=49289112#49290012 . - that turned out to be my fault, not the model. My SVG rendering software was stripping some attributes. Here's the Gemini 3.7 Flash pelicans in the fixed renderer: https://tools.simonwillison.net/markdown-svg-renderer.html#u...

>> Most importantly, the pelican has one leg on each side of the bicycle - that's very rare. How can you say "each side"? I don't see any z-ordering between bicycle frame and legs.

Because, if you look at the image you can very clearly see that the frame and relative components partially occlude one leg and not the other.

The model is presumably aiming for an acceptable visual representation, not trying to produce z-ordered components and a lost of any other random requirement people might come up with. The task is to show a pelican riding a bike, not produce a technical design that is layer order correct, after all.

Also:

> Near leg: from (300,245) to (356,408): thigh+shin as a single slightly bent line: M300,245 C 310,320 330,370 354,406. Stroke #F2953F width 12, linecap round. Far leg: from (320,250) to (404,452): M320,250 C 350,330 385,410 402,448. Stroke slightly darker #E0802F (behind, drawn before near leg but after bike? Pelican legs are in front of frame? Pelican is drawn after bike, so legs overlap the frame. The far leg ideally would be behind the frame, but in flat cartoon this is acceptable — or draw the far leg before the pelican body but after bike; overlaps the red frame.

Re: Qwen 3.8 27B

#749
post #615
post #550

Earlier quoted context omitted.

I was quite impressed by Muse Glimmer, and while I am sure people will observe that it is less good on benchmarks, my first experiences with this new 27B have been somewhat exasperating, whereas testing Muse Glimmer was rather fun. I have not tested either in an agentic context, mind you.

Glimmer works really well as an "explore" agent model (like in Opencode.) It seems to be extremely efficient at searching and collating that info, and executing commands. From my testing so far, Qwen 3.8 is better at code but it tends to meander and take forever if it has to look in a lot of places. Glimmer will use like ~1k tokens to formulate a plan and Qwen 3.8 will routinely go over 10k

Have you tried turning down the new Qwen's reasoning effort level from xhigh, which it defaults at?

LM Studio isn't exposing a dropdown for this, at least with the unsloth build.

Unsloth Studio / Desktop does.

Re: Qwen 3.8 27B

#750
post #371

Earlier quoted context omitted.

I absolutely love this comment. I wished there was a website where people would post their working command lines as well as what hardware they are using to run that stuff on + tokens / sec prefill + gen.

pretty sure this exists already...

Feel free to link it if it does…
Post reply on HN