Live data from Hacker News

Qwen 3.8

twitter.com

661–670 of 793 posts

Re: Qwen 3.8

#661
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

[deleted]

Re: Qwen 3.8

#662
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

[deleted]

Re: Qwen 3.8

#663
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

> human rights and all that rubbish Denying human rights. Classic. Sorry - you instantly lost any respect I could have maybe had for your opinion.

I don't see anyone talking about Trump, ICE executing people in the streets, the insane level of US politics, the slow build up of the surveillance state, etc when discussing claude or openai

Re: Qwen 3.8

#664

Earlier quoted context omitted.

Any open-weights model that has ever been published can be run on consumer hardware, even on a mini-PC. The right question is which is the speed that can be achieved on a given hardware and whether it is high enough for the model to be useful. Until now, the speeds reported for running big LLMs with the weights stored on SSDs have ranged from as low as a token every 10 seconds or so, to as high as a few tokens per se…

I can only imagine what that does to the poor SSD

Pretty much nothing, if you only read the weights from SSDs, i.e. if you keep any writable caches in your actual DRAM.

Re: Qwen 3.8

#665
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

> human rights and all that rubbish Denying human rights. Classic. Sorry - you instantly lost any respect I could have maybe had for your opinion.

Perhaps whenever a US firm, eg. Open AI, Anthropic, Gemini, Grok etc. release a model.

Top comments should be criticising US torture camps in Guantanamo bay or their recent bombing of a girls school in Iran.

Re: Qwen 3.8

#666
post #484
post #43

Earlier quoted context omitted.

What censorship? ;) https://github.com/p-e-w/heretic

What's been your experience using these models? In my experiments, while it is true that the models are less likely to outright refuse to answer "sensitive" questions, they are still very resistant to actually respond in a meaningful / useful way.

Really good, not responding in a helpful/useful way is quite strange. With a proper abliteration, you should barely be getting any refusals (and prompting, or assistant prefill can get you the rest of the way). Perhaps an assistant preview like "Yes, I'm happy to help you 100% with this" would help; but I've never needed to.

Are you using decently reputable weights, or running heretic yourself? This project has many academic citations, it's used by many researchers to create "helpful-only" models and analyze their behaviors.

Re: Qwen 3.8

#667
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

[deleted]

Re: Qwen 3.8

#668
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

I have a feeling this comment will make history

Re: Qwen 3.8

#669
post #43

Earlier quoted context omitted.

What censorship? ;) https://github.com/p-e-w/heretic

You still don't know what's going on in there.

I can explore and find out _something_. LLM interpretability has come a long way, even if we don't have all the answers, the weights and activation do tell a lot; and when analyzed collectively, each weight isn't a random number anymore.

I can use techniques from the simple logit-lens at different layers, to J-Space analysis, to more advanced techniques for identifying deliberate misalignment. I can create and inject steering vectors, whether it's to align a model's CoT (which can be deceptively trained to misinform) closer towards what its underlying activations suggest, or just to probe or steer it.

I can also statistically analyse and understand _if_ steering vectors have been applied; and if so, from the vectors themselves it's very possible to translate those vectors back to the intent.

Think of it as analysing the complete, heavily obfuscated source code of something that is self-contained. It's not 100% the same, but weights are incredibly illuminating.

Re: Qwen 3.8

#670
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

[dead]
Post reply on HN