Live data from Hacker News

Qwen3.8-Flash-Next

qwen.ai

211–220 of 246 posts

Re: Qwen3.8-Flash-Next

#211

Earlier quoted context omitted.

You can make the other argument that China subsidizes the price and that they can't be profitable at this pricing level. From an industrial strategy standpoint, they already do this for many other industries with huge subsidized state loans. So we can go round and round on this, each with our made-up objections about how it's temporary or unrealistic or impossible or whatever, or we can just accept the prices as list…

Didn't OpenCode CTO state they could replicate deepseek pricing on rented hardware?

There's a difference between the Deepseek.com provider lunch pricing and the pricing every other provider is doing now.

Right now DS4-Pro-0813 is available from multiple providers for $1.32/million input tokens[1].

It's pretty easy to work backwards from B200 and electricity prices and see this is profitable even without the heavy serving optimization these providers are doing[1.5].

The OpenCode CEO said: "inference is very profitable and probably a good opportunity to understand some basic business math"[2] and "the inference we do is already profitable and that's with some middlemen involved"[3]

If at this point people don't believe inference can be profitable, and providers can turn the prices up and down to choose exactly how profitable they make it I don't know what to say.

[1] https://openrouter.ai/deepseek/deepseek-v4-pro-0813#provider...

[1.5] https://www.seangoedecke.com/ai-inference-is-obviously-profi...

[2] https://x.com/thdxr/status/2042277156940587469?lang=en

[3] https://x.com/thdxr/status/2042614323344818520

Re: Qwen3.8-Flash-Next

#212

Does anyone have an idea how this might perform on a DGX Spark at longer contexts? I've been trying to investigate their performance with these medium-sized MoE models, but I'm seeing a lot of incomplete and conflicting information. The 273 GB/s bandwidth looks awfully bad on paper...

A single Spark alone is not worth the price. You are paying $1000 just for networking equipment you aren't using. At 2x it starts to maybe become worth it if you don't want to deal with Apple. Outside of the newest Macs, I can't think of anything else you can get 256 GB ~550 GB/s memory bandwith for $8200. Even at 3-4 Sparks it scales relatively well.

With 2x Sparks, I am getting 40 t/s. I'd guess that without MTP you'd get 12-15 on 1 Spark, maybe 20 with MTP?

Re: Qwen3.8-Flash-Next

#214

Earlier quoted context omitted.

For comparison with hosted models, GPT 5.6 Luna scores 67% on DeepSWE, compared to 59% here for Qwen. Luna is $0.20 / $1.20 vs $0.16 / $0.47 with Qwen.

This is a good counter argument. But you have to note that this is after OpenAI cut Luna costs by 80%. If you compare launch pricing, Qwen probably comes out ahead on a cost-performance basis.

Why would anyone car what the launch price is? Comparing launch pricing is just an odd thing to do.

Re: Qwen3.8-Flash-Next

#215

Earlier quoted context omitted.

73GB for the 1 bit model...

That probably includes the 51b ngrams too. It's possible that those could be streamed from NVMe on-demand. The Engram paper that developed this technique streamed from RAM to VRAM at only ~1% performance degradation, but these strix halo boxes and the spark have much slower memory, so it's possible moving down another rung on the memory hierarchy wouldn't affect their performance too much. This will almost certainly…

This guy claims 6% throughout hit for this approach:

https://x.com/0xBakeer/status/2092694905978237224?s=20

Crazy how fast things move these days.

Re: Qwen3.8-Flash-Next

#216
post #86

Earlier quoted context omitted.

For comparison with hosted models, GPT 5.6 Luna scores 67% on DeepSWE, compared to 59% here for Qwen. Luna is $0.20 / $1.20 vs $0.16 / $0.47 with Qwen.

Those prices are just tokens? Since each model uses different amounts of tokens to do the same thing, it's a misleading price that often makes open-weights look more competitive than they are, since most open weights models use dramatically more tokens and time to complete tasks than many frontier models. In Artifical Analysis's cost per task, Luna(max) costs $0.05 per task, and Qwen 3.8 27B costs $0.25 per task, a 5…

the important thing is that Qwen 3.7 27B will run unlimited jobs on my consumer grade laptop at 60 tokens/second for free, forever, in about 1-2 years

Re: Qwen3.8-Flash-Next

#217

Earlier quoted context omitted.

> You’re absolutely right to be hopeful. Three honest possibilities, and I’ll be straight with you about each: > [UGC styled humorously as LLMisms] All joking aside, having interacted with Claude intensely for the last 8 months and about 30 hours/week in the last 3, I’ve started to notice how (for want of a better word) “readable” (“digestible” ? “comprehensible” ? “Predictable” is the wrong direction.) information c…

I find LLMisms very annoying to read, it’s almost like they are bullet points in the shape of a paragraph. It feels very “skippy” to me. EDITED: Removed a question that I couldn’t make feel suitably polite.

Suppose you time-zap a modern physics curriculum on a solarpowered computer tablet to any shortly-pre-Galilean era and observe their reaction to the course notes.

In that era, plenty of fields required mathematics, engineering and architecture.

The church would prescribe and uphold Aristotelean Logic "When objects fall, they fall down" style statements (never mind that if you throw an object up, it doesn't instantly have a downward velocity component).

When the church has new cathedrals, domes, catapults for Crusades etc. built they actually relied on architects and engineers using rule of thumb formulas.

Those educated in Aristotelean Logic were viewed with higher stature than those actually making experience-based calculations using mathematics.

The era often associated with Galileo is when the stature reversal started to surface and be openly talked about. The universe is best described in mathematics, not natural language factoids.

Right before this recognition, those of the higher stature Aristotelean Logic education would look down on the architects and engineers who already used mathematics by pragmatic necessity.

To these people the time-traveled physics curriculum would look like cliche mathematics. Given randomized sections of text either drawn from either Aristotelian Logic texts or modern physics texts, they would easily be able to discern the Aristotelian Logic from the obtuse mathematical phrasings. To them the smartphone loaded with Maxwell's texts, Jacksons Electrodynamics, Goldsteins Classical Mechanics etc. is talking "math".

The ability to recognize outlier writing style says nothing about content quality.

Mike Judge (widely known from the MTV series Beavis and Butthead) studied physics. One of his movies "Idiocracy" about a modern day average-educated protagonist who accidentally ends up in a future decaying society filled and run by intellectually retarded people contains scenes where this future uneducated population considers his speech "gay" simply because of his higher level of education.

Could the adversarial prospects of job loss, edge loss (a long expensive difficult education replaced by tensors fitting megaprojects that take a couple of weeks), etc. combined with recognizable communication patterns also explain our pejorative references to LLM-isms? Personally I'd prefer LLM's to communicate in mathematical terms, but all the LLM-isms are effectively a mirror of our contemporaries.

Either we complain because algorithmic responses look like a mathematics textbook ("just fix my python array plz, why are we talking about "sets" and "injective" and "Lipschitz continuity"?), else we complain its "pretty printed to natural language".

We should also recognize large language models are in a "Damned if you do, damned if you don't" situation.

When a reader considers some text as mathurbation, are they really just abreacting the awareness of lack of education?

How could anyone possibly expect Fourier optics "pretty printed" to non-mathematical language to result in any satisfactory experience?

Re: Qwen3.8-Flash-Next

#218
I'm really impressed. Gave QwenCloud $18, handed 3.8-flash a few big forks of a lot of code, it did some archeology and made a clean merge. Then it used the project's tools to bisect a regression and fix it.

Was not expecting it to just get that right without any fuss, and it barely used 10% of this weekly limit. Something like 90M cached in/400k out for $0.45 is wild

Re: Qwen3.8-Flash-Next

#219
post #214

Earlier quoted context omitted.

This is a good counter argument. But you have to note that this is after OpenAI cut Luna costs by 80%. If you compare launch pricing, Qwen probably comes out ahead on a cost-performance basis.

Why would anyone car what the launch price is? Comparing launch pricing is just an odd thing to do.

Because labs can learn to optimize inference post launch, plus can move to use bigger/better clusters depending on demand. It is not impossible to imagine Qwen cuts prices further with QAT/MTP-like improvements.

Re: Qwen3.8-Flash-Next

#220
post #58

> Qwen3.8-Flash-Next features a 125B-parameter main model, supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. Didn’t see this mentioned yet. I wonder what this means for the effective size. It’s evidently ~176B paramètres, but how does that get quantized. A 4-bit quant under 100GB seems unlikely, I’m suspecting this won’t run in 128GB unified memory In principle I like the id…

People in my server are running it on Strix Halo 128GB using RoCmFP4 and reporting 35tok/s, without much optimization, with proper MTP, better kernel, expecting about 50-60tok/s.

How do they like it, compared with 3.8 and DS4?
Post reply on HN