This one should work on personal computers! I'm thankful for Chinese companies raising the floor.
[flagged]
Astroturf/shill accusations are also against the HN ethos. https://news.ycombinator.com/item?id=11257034
21–30 of 64 posts
This one should work on personal computers! I'm thankful for Chinese companies raising the floor.
[flagged]
Astroturf/shill accusations are also against the HN ethos. https://news.ycombinator.com/item?id=11257034
Earlier quoted context omitted.
https://artificialanalysis.ai/leaderboards/models?open_weigh...
Compare these rankings to actual usage: https://openrouter.ai/rankings Claude is not cheap, why is it far and away the most popular if it's not top 10 in performance? Qwen3 235b ranks highest on these benchmarks among open models, but I have never met someone who prefers its output over Deepseek R1. It's extremely wordy and often gets caught in thought loops. My interpretation is that the models at the top of Artific…
Earlier quoted context omitted.
[flagged]
Let's say any country create the most powerful - and thus best - LLMs. They over time infiltrate it with their political will. Over 20-30 years, I'd imagine people asking those LLMs will have their minds' shifted. But. That's just me, my pessimism-sci-fi scenario.
Earlier quoted context omitted.
https://artificialanalysis.ai/leaderboards/models?open_weigh...
Compare these rankings to actual usage: https://openrouter.ai/rankings Claude is not cheap, why is it far and away the most popular if it's not top 10 in performance? Qwen3 235b ranks highest on these benchmarks among open models, but I have never met someone who prefers its output over Deepseek R1. It's extremely wordy and often gets caught in thought loops. My interpretation is that the models at the top of Artific…
of course, this is a politically charged subject now so fair assessments might be hard to come by - as evidenced by the downvotes i've already gotten on this comment
Is there a crowd-sourced sentiment score for models? I know all these scores are juiced like crazy. I stopped taking them at face value months ago. What I want to know is if other folks out there actually use them or if they are unreliable.
Earlier quoted context omitted.
Compare these rankings to actual usage: https://openrouter.ai/rankings Claude is not cheap, why is it far and away the most popular if it's not top 10 in performance? Qwen3 235b ranks highest on these benchmarks among open models, but I have never met someone who prefers its output over Deepseek R1. It's extremely wordy and often gets caught in thought loops. My interpretation is that the models at the top of Artific…
Thanks for sharing that. Interesting that the leaderboard is dominated by Anthropic, Google and DeepSeek. Openai doesn't even register.
Earlier quoted context omitted.
[flagged]
I’m American. Just giving some background to the feeling. There’s some discontent with some western communities (localllama) that Chinese developers have been open weighting all of their models while most western models have been closed weights.
Right now though China is dropping huge improvements across the entire spectrum of model sizes with Qwen, Kimi, DeepSeek, GLM, and Yi. We've also got Mistral doing competitive self-hosted models too, but they're French. Local AI tooling is plainly _not_ being driven forwards by the United States.
Is there a crowd-sourced sentiment score for models? I know all these scores are juiced like crazy. I stopped taking them at face value months ago. What I want to know is if other folks out there actually use them or if they are unreliable.
openrouter usage stats
It's an interesting proxy, but idk how reliable it'd be.
Earlier quoted context omitted.
I’m American. Just giving some background to the feeling. There’s some discontent with some western communities (localllama) that Chinese developers have been open weighting all of their models while most western models have been closed weights.
Meta seems to have now stepped out of the running despite being the local LLM catalyst. Anthropic has done nothing. IBM's Granite and Microsoft's Phi are both very far behind. AWS doesn't even attempt to compete. Grok failed to make good on their promise. OpenAI only even entered the game yesterday so it's hard to tell if they're actually serious since they released such an overly censored model that isn't really bet…
This is why the Chinese labs are so open, they don't ever need to make a profit, they just need to make good AI.
Earlier quoted context omitted.
[flagged]
Let's say any country create the most powerful - and thus best - LLMs. They over time infiltrate it with their political will. Over 20-30 years, I'd imagine people asking those LLMs will have their minds' shifted. But. That's just me, my pessimism-sci-fi scenario.
But still, the most recent version of american foss model gpt-oss is just so filled with censorship that its just not worth it in the name of "safety", so to me both are doing censorship but I'd much rather use chinese censorship since its only censored on chinese topics and I mean, I personally wouldn't be ever asking chinese models chinese questions but maybe that's just me.
And even if I would, I would probably ask it on some uncensored, in fact I was actually thinking of creating a fine tune like the perplexity, or just this idea to break chinese censorship.
I also think that some better idea needs to come up with multi modal approach so that censorship could be removed by mixing and matching american and chinese models, I do think it is far from reality but I read a recent comment about harmony which gpt-oss uses and it does look promising I am not sure.