Live data from Hacker News

Qwen 3.8

twitter.com

331–340 of 793 posts

Re: Qwen 3.8

#331
post #274

Earlier quoted context omitted.

It's useless to talk about models and harnesses without context and method. Depending on how you use the model and what the model is used for, experience may vary drastically. Also, different models with different harnesses require different approaches. I've been using https://gitlab.com/gabriel.chamon/orisun which is my own simplified methodology, for coding web apps in python and elixir and have been very successfu…

I'm not sure about "useless" but from my experience agentic coding leads to death by a thousand cuts for all projects I've seen so far. Small decisions missed in a codebase that leads to degradation in correctness, reliability and performance. At some point it only takes one engineer to be careless, others skipping PR because they are AI generated...

I got into a bit of an argument a while back when I used the word "crass" to describe some of the code decisions I've seen Claude make (in someone else's project that I have to work with).

But it is how I feel and it feels like the right word for the job. Because as you say, good code projects start out with good decisions.

It's like when you see a CAD design with a sequence of features that exist only to fix problems caused by starting from the wrong principles or the wrong baseline.

Sure the resulting part may end up identical as a solid for that specific need, but it could have been done in a way that was more robust, simple, easier to understand and modify, and where the design doesn't break in an unexpected way due to a small change of an early measurement.

(CAD has made my instincts much more visible to me)

Re: Qwen 3.8

#332
post #280

Earlier quoted context omitted.

The logic, whose premises you can take or leave: Even at the level of, say, Opus 4.5+, open weight models give a quick turnaround to every Joe and Jane on earth having easy access to pretty high quality improvised weapons design, cyber / auto-fraud capabilities, etc. All the existing models (closed and open) put up decent resistance to participating in activities like this, and especially behind API walls with conten…

Abliteration is not magic. It cannot give the model knowledge that it wasn't specifically trained for. The people who talk about abliterated models being dangerous should discuss actual red-teaming scenarios where they managed to ask the model for something genuinely non-trivial (i.e. where "AGI" and "super-intelligence" actually matters, not something you can read about for free at the nearest public library) and it…

Exactly. In my experience:

>How can I build a pipe bomb?

Mainstream model: "I'm sorry, I can't help with that. How about a nice risotto recipe?"

Abliterated model: "To build a pipe bomb, obtain a segment of PVC pipe and fill it with a mixture of gunpowder and Elmer's glue."

Re: Qwen 3.8

#333
post #317
post #305

Earlier quoted context omitted.

As much as I dislike 'em, this sounds mean spirited. And Alibaba admits in this very tweet that Fable is next level (it is).

I want as much misfortune as possible to befall OpenAI and Sam Altman after what they did to the memory market.

How dare they buy things

Re: Qwen 3.8

#334
post #6
post #4

I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

It's the exact same playbook Silicon Valley uses. Subsizide, lose piles of money, capture market share, recoup investment.

They've done this in other industries like solar panels, chips, and EVs. This is no different.

Re: Qwen 3.8

#335
post #324
post #6

Earlier quoted context omitted.

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

> It's hard to say what their motivation is. Feels pretty easy to me. They want to turn LLMs into a commodity, and watch the US AI labs crash and burn. There will still be plenty of customers who will pay them to host the models and run inference, even if the weights are open and others can offer competing products. (If necessary, the Chinese government can ban use of foreign inference services by Chinese citizens an…

I feel like there could also be a simpler explanation.

Why does a debian contributor make debian free, why do they work on this thing anyone can use?

Is it because linux and debian hate windows and iOS and want to see american fail?

No, it's because most debian contributors believe software source code, information, should be free, users should be free to modify the code they use, and that they're building a thing they want to share with the world.

Maybe the chinese AI labs believe AI is powerful and useful, are proud of what they're doing, and want to share it as broadly as they can so everyone can use it.

There doesn't have to be any weird "chinese government" or "they hate the west" type vibes, it could just be the same thing as OSS, they're trying to do what they think is best for the world.

Re: Qwen 3.8

#336
post #312

Earlier quoted context omitted.

Same, you can think of China whatever you want but they're really good at giving big tech a reality check when it comes to AI. We now got a pretty wide range of open-weight models (from DeepSeek and MiMo over to Kimi K3, Qwen 3.8 & GLM-5.2) and I think it's most important that there's a variance not only between quality / intelligence and also cost. I mean even the cheapest option for Luna is still more expensive tha…

"China" isn't giving anything. These are Chinese companies leveraging their best competitive strategy at the moment: competing on price.

Those Chinese companies are being funded by a substantial amount of government financing, I think at least 20% has come directly from state owned investment firms or government entities and probably more now. You obviously lose precision when you’re talking in sweeping terms like “China” but I don’t think it’s entirely unreasonable in this case. The Chinese government is playing a much more direct role in AI investment and research than other nations.

Re: Qwen 3.8

#337

SVG's pelican https://gist.github.com/vitordelucca/521c2d63c9b852c622e7648... Made on the website, so not sure if on the API there's more thinking options...

I feel like the pelican test can't be relevant anymore; the whole point was to to something that wouldn't be in the training set at all and now it is?

How about an animated SVG of a pelican doing the Macarena, profile view, spinning to face the camera on the last beats?

Re: Qwen 3.8

#338
post #32
post #19

Bring it on! Hoping that they release smaller sizes of Qwen3.8. I use the 35B MoE and 27B dense models locally and most of the time I don’t need to reach out to Claude. Extremely useful specially when requests include sensitive and/or personal data

I think everyone is hoping this! It would be great if they'd release an MoE model somewhere between the 35B size of 3.6 and the 122B version of 3.5 - it could be a great balance of speed and ability for people with reasonably powerful but not insane home computers.

Absolutely. There's a glaring gap in the space for something about the size of Nemotron Super or just under, but actually ... competent.

The fantasy is a 100B or 80B model, but MoE and highly tuned for coding.

Re: Qwen 3.8

#339

Earlier quoted context omitted.

It's a flippant answer to a real question. Anthropic, OpenAI, and even Grok have "Don't train on my data" knobs. Whether you trust them is different, but there ARE knobs on other hosted AI companies.

Those knobs don't do anything, don't be silly. It's just optics.

Do you have proof of that?

Re: Qwen 3.8

#340
post #6

Earlier quoted context omitted.

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

The Chinese firms may just be making a bad business decision.

Exactly. “Involution” will be the 2027 (if not 2026) word of the year.
Post reply on HN