Live data from Hacker News

Show HN: Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh

huggingface.co

11–14 of 14 posts

Re: Show HN: Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh

#11
post #2

I will TLDR you on our thought process, research, training and benchmarks. 1. When running our quantized Qwen 3.8 27B instances we were very annoyed by random reasoning loops (in the paper bellow refered to as "overthinking errors". These random loops were persistent throughout medium and low reasoning settings. 2. We found a paper by Meta that's supposed to target this phenomenon in PTQ, but when used straight out o…

Nice. Going to dl and give it a whirl.

Re: Show HN: Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh

#13
post #2

I will TLDR you on our thought process, research, training and benchmarks. 1. When running our quantized Qwen 3.8 27B instances we were very annoyed by random reasoning loops (in the paper bellow refered to as "overthinking errors". These random loops were persistent throughout medium and low reasoning settings. 2. We found a paper by Meta that's supposed to target this phenomenon in PTQ, but when used straight out o…

I tried an oMLX quant of your model (suzu89/Swift-Qwen3.8-27b-oQ8-mtp -- not mine) and liked it. It certainly seems to cut down on thinking compared to stock Qwen3.8 27B in my (limited) testing.

A couple of questions:

- Have you tried the peculiar-ragdoll/Qwen-Sharp-Chat-Templates with it? They replace the default chat_template.jinja with one that encourages less thinking.

- When are you releasing Swift-Qwen3.8-Flash-Next? :)

Post reply on HN