Live data from Hacker News

Qwen3.8-Flash-Next

qwen.ai

41–50 of 246 posts

Re: Qwen3.8-Flash-Next

#41
Pelican: https://gist.github.com/SerJaimeLannister/8fdef9c00175da0ca6...

Aside from the pelican, I am sort of impressed by the fact that things are going the way in terms of really impressive small models.

Also I love how this uses N-gram embedding. I think that Longcat was the first one who used it (I submitted that submission on hackernews because I really just loved the idea of it that I understood), I am certainly more interested in local LLM models and its interesting how they are utilizing new architectures to do some really impressive optimizations!

(Do note that I created it using a free rate limited end-point that I found on the huggingface space section: https://victor-chat-with-qwen3-8-flash-next.hf.space)

Re: Qwen3.8-Flash-Next

#43
post #12

Earlier quoted context omitted.

I don't like these comparisons. Sure it is impressive, but it does not have a world knowledge of larger models. It has most of theirs intelligence.

For world knowledge, you'd want it to find and reference the source material to be sure. At that point, it doesn't matter if the knowledge is embedded.

Keep in mind a web search might not include scanned books baked in the weights ;)

Re: Qwen3.8-Flash-Next

#44
post #39

Earlier quoted context omitted.

This is a good counter argument. But you have to note that this is after OpenAI cut Luna costs by 80%. If you compare launch pricing, Qwen probably comes out ahead on a cost-performance basis.

The luna cost cuts were real though, not a one time promotion or something, due to some optimization (probably distillation?) that openai did.

Was it?

Given the timing, I think they A. shat their pants since Deepseek flash just came out with insane pricing before the price hikes, and B. Anthropic is really struggling in model tiers below opus.

It was smart for them to cut prices regardless of whether they had 80% efficiency gains or not

Re: Qwen3.8-Flash-Next

#45
post #34

It looks like this also undercuts the already absurdly inexpensive Deepseek Flash in pricing. Wild.

Where are you seeing that? At the bottom of this post from Qwen I see: Qwen 3.8 flash: $0.16 / $0.47 Compared to Deepseek 0723: $0.03 / $0.075 (units in USD/m tok)

Deepseek 0732 is $0.22/$0.66 off peak

Re: Qwen3.8-Flash-Next

#46
post #12

Earlier quoted context omitted.

I don't like these comparisons. Sure it is impressive, but it does not have a world knowledge of larger models. It has most of theirs intelligence.

For world knowledge, you'd want it to find and reference the source material to be sure. At that point, it doesn't matter if the knowledge is embedded.

I think the big models have adequate recall, so tool use is probably unnecessary, but the user said the correctness of my response is important. Let me look up the data instead of relying on my memory.

Re: Qwen3.8-Flash-Next

#47
post #12

Earlier quoted context omitted.

I don't like these comparisons. Sure it is impressive, but it does not have a world knowledge of larger models. It has most of theirs intelligence.

For world knowledge, you'd want it to find and reference the source material to be sure. At that point, it doesn't matter if the knowledge is embedded.

World knowledge also means knowing the various algorithms and ways particular programming problems are solved.

You can't search what you don't even know exists.

Re: Qwen3.8-Flash-Next

#49
post #25

It's in Unsloth Desktop already. Looks like it's 73GB, so 128GB Mac or Strix Halo etc will work. Exciting!

I only see a 1-bit quant posted on unsloth HF and it’s 72.5 GB. Is that what you mean? That’s much bigger than I expected. If you can’t run a 4 bit quant in on Strix Halo it becomes a lot less interesting. https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF
Post reply on HN