Live data from Hacker News

Qwen 3.8

twitter.com

131–140 of 793 posts

Re: Qwen 3.8

#131
post #90

Earlier quoted context omitted.

DeepSeek V4 pricing is insane, 10x-30x cheaper to use than most other models, and it usually is good enough for most tasks.

It doesn't matter if it's cheaper, specially if it consumes more resources to do the same task as the competition Besides, in a few days, they'll change their pricing, doubling it during their peak hours, so, realistically: - It will be 2x more expensive if you live in their time zone - It will be 1.5x more expensive if you live in a time zone that is adjacent to theirs - It will be the same price IF you use it while…

I was not referring to the input/output price, but the cost of doing a specific tasks, in practice it is ~10x cheaper than GLM-5.2 for example, to accomplish the same task (for the tasks it can do).

I have been happily using DeepSeek V4 Flash for the last couple of months now. I tried GLM-5.2 for a while, but it was too slow and verbose compare to DeepSeek V4 Flash. If I have a basic skill I need to execute, DeepSeek V4 flash is still the best model for it.

Re: Qwen 3.8

#132
post #19

Bring it on! Hoping that they release smaller sizes of Qwen3.8. I use the 35B MoE and 27B dense models locally and most of the time I don’t need to reach out to Claude. Extremely useful specially when requests include sensitive and/or personal data

That's a 2.4T model, how would they reduce this to 35B and still give some accuracy? That's a completely different arch.

there's been a lot of research about reducing models by taking out layers; there's also using it to train smaller models by optimizing parameters.

I dont see most model building as anything more than a pig at a slop troth, despite the level of sophistication; they're still rarely pruning the input beyond random sampling.

Re: Qwen 3.8

#134

Earlier quoted context omitted.

Who do you buy DeepSeek from? I bought it through OpenRouter and used it with Pi agent. The model was good, but there appeared to be a pricing glitch or something, because it burned through $50 in under an hour on pretty trivial stuff. Pi agent claimed it only used like $1. OpenRouter claimed differently and said I used all $50.

Check cache hits in your logs. You can use Openrouter or pi config to pin providers with best cache hit rates (or disable ones with the worst). I use Openrouter for everything except Deepseek. For Deepseek I use their API directly.

There is a 3rd party harness specifically tuned for deepseek (reasonix). Have you tried that?

Re: Qwen 3.8

#135
post #6

Earlier quoted context omitted.

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

Humanity is a bit of a stretch, and to be seen over time, not that I'm saying it won't happen; let's get some hubris here.

The whole discussion is hubris. This is a discussion of a twitter post about something that is announced to happen but hasn't yet.

Not one person here has any idea what is going to happen long term.

Re: Qwen 3.8

#136

I predict that no one will use this and everyone will use Kimi K3.

I've been playing around with K3 a bunch, but the verbosity of the reasoning makes complete e2e agent work basically cost the same as other smaller models, and I'm not seeing a huge difference in quality, just a way longer e2e completion time.

Re: Qwen 3.8

#137

Earlier quoted context omitted.

> Not that hard to say IMO, Unless you work there, your opinions are guesses, and parent is saying we cannot know, which remains true even with your guesses :)

Same can be said for every companies decisions then. Why does Antropic not open source their best models? My ”guess” is it’s because they are printing money with their closed models

Right and what other product does Anthropic really have besides?

Re: Qwen 3.8

#138
post #123
post #109

Earlier quoted context omitted.

Is qwen 3.6 27b the best model you can run locally at the moment? Not that I have the VRAM for it, but just curious.

In my experience, yes. A bit more reliable than gemma for me. I mostly use A3B (35B, mix of experts) though, because it‘s faster, and in the same ballpark intelligence wise as the dense 27B, so it’s the sweetspot for me. I want to try cohere‘s mini code model next, but worried the runtimes aren‘t optimized for that yet.

I found qwen3.6:26b slightly better on my 32G mac mini than the same sized gemma until gemma was updated with better tool support 4 or 5 days ago.

It is like a ping-pong game: the advantage flips back and forth between providers.

Re: Qwen 3.8

#140
post #78
post #6

Earlier quoted context omitted.

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

There’s a Twitter thread making rounds by Dean Ball about deceleration in AI development caused by open models and I can’t understand how people don’t see that it’s true: open models dismantle the frontier lab capex spend potential by reducing the training budget to zero in the limit. Tokens from different providers are not fungible, but customers are nevertheless very price sensitive and close enough is good enough,…

> there’s a chance frontier models get sponsored by the USG and are never released publicly so they can’t be distilled and then what?

That premise hinges on one implicit assumption: Chinese advances are due to distillation ONLY and that Chinese model providers cannot keep advancing if they do not distill, which is a very big if. If Chinese models keep advancing in such a scenario, and they almost certainly will, they will overtake publically available models by US providers and China will dominate the LLM industry.

Post reply on HN