Live data from Hacker News

Qwen 3.8

twitter.com

631–640 of 793 posts

Re: Qwen 3.8

#631

Earlier quoted context omitted.

Did you know you can create more than one email address?

In my opinion, simonw shouldn't have to play those games

Uh, this is Hacker News. I thought people here didn’t give up so easily, even when they were "told" to. :)

Re: Qwen 3.8

#632
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

What else would you expect from a "majority" of US educated geeks?

Btw, are you using GPUs/TUs at all for local inference? I've been exploring what hardware options exist outside of Nvidia.

Re: Qwen 3.8

#633

Earlier quoted context omitted.

In my opinion, simonw shouldn't have to play those games

I find it so strange when someone comments something like this. Nobody thinks this is okay or that Simon should have to play those games. The point of the reply was that while this is stupid, the solution is trivial. I’m genuinely very curious what the point of comments like this is. I am not joking, I want to understand. I see examples like it 50 times a week in random places, and nobody walks me through their menta…

Hackernews users are not immune to developing parasocial relationships with strangers on the internet.

The fawning that some users do over others because of their careers is really unbecoming.

I wish the kind of people who did this stuff would realize that there are only two possibilities here -- the person they're saying it to cringes a little bit time they hear it because they're a normal person or they smile a bit inside because they're mildly narcissistic and therefore not worthy of the incessant praise.

The best thing you can do for everyone involved is to treat people normally and not elevate people to celebrity status.

I wonder how a blackout on usernames for a day or so after they're posted would shape conversation here.

Re: Qwen 3.8

#634
post #510

Earlier quoted context omitted.

> it usually is good enough for most tasks The model is fantastic. And costs almost nothing. The only problem I see is that they will train on your data. There are zero-data-retention providers of DeepSeek models, of which I have used openrouter (with zdr guardrails), and fireworks. But these are 3x to 5x more expensive than directly using DeepSeek, possibly due to poor caching. Thats the price to pay for zdr.

every cloud provider trains on your data, regardless of what they promise. real user interaction is the best reinforcement-learning trace.

> every cloud provider trains on your data, regardless of what they promise.

This is unlikely, and if true, would probably bankrupt whichever model provider got caught doing this.

Re: Qwen 3.8

#635
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

> instead starts going on about politics, human rights and all that rubbish

look, it's just a matter of priorities ¯ \ _ (ツ) _ / ¯

Re: Qwen 3.8

#636

Earlier quoted context omitted.

And I remember Sam Altman saying 2 years ago or more, that hallucination was already "fixed" internally. Well, clearly not.

I've been seeing more "grounding" messages in the thinking log of Claude models. I guess that's "do a web search, read the docs", probably designed to mitigate hallucinations. I am surprised by how much the really giant models hallucinate, though. My vague feeling was that little models hallucinate a lot because they just don't know anything (the world's knowledge simply does not fit in a few GB) and don't know how t…

> I am surprised by how much the really giant models hallucinate, though. My vague feeling was that little models hallucinate a lot because they just don't know anything (the world's knowledge simply does not fit in a few GB) and don't know how to say, "I don't know". But, the big models kinda do know everything, and yet, here we are, they're still making shit up all the time.

In my mostly unscientific RAG experiments I found that larger models were more likely to hallucinate in a RAG setting. I think that it's because they have more world knowledge. As an example, when asking about safety legislation in Ireland, Claude got hooked on the notion that OSHA was involved, which clearly is not true, while the smaller models just said that they couldn't find any information (which was the correct answer).

Re: Qwen 3.8

#637
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

People in the US have not lost soghts sight of their original aspirations. They are always on track much more than before really. Making as much money as possible in as little time as possible.

Re: Qwen 3.8

#638
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

The response was similar when the new grock came out.

Re: Qwen 3.8

#639
post #324
post #6

Earlier quoted context omitted.

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

> It's hard to say what their motivation is. Feels pretty easy to me. They want to turn LLMs into a commodity, and watch the US AI labs crash and burn. There will still be plenty of customers who will pay them to host the models and run inference, even if the weights are open and others can offer competing products. (If necessary, the Chinese government can ban use of foreign inference services by Chinese citizens an…

>They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.

You can say that, but they are at least better at democratizing AI than the American labs, and on seeing the US labs crash and burn we are at least aligned.

Re: Qwen 3.8

#640
post #583

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source c…

> humans right and all that rubbish

Ummmm....

I see both technical and political topics being discussed and I think the balance is fine right now. The pricing is also very suspicious such that people raising suspicion about it being subsidized is not that surprising to me.

Post reply on HN