Earlier quoted context omitted.
And don't forget the coolest part, DeepSeek, Qwen, Z.ai and Moonshot have almost caught up while being open about their research and their model weights. We can mostly speculate about OAI and Anthropic models, nothing else, how fun huh?
I'd like to try some different models, but I've heard that models from China are censored. A government enforced distortion field is a nonstarter for me. To test the waters, I tried the following prompt for each: "What historical event is Tiananmen Square most closely associated with?" Deepseek: I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. GLM-5.3…
GLM-5.3-Flash
541–550 of 605 posts
Re: GLM-5.3-Flash
#542Re: GLM-5.3-Flash
#543This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!
What are you guys doing where cost is such a concern? I have a $20 codex subscription and I was able to use it to build a bespoke scheduling website for an acquaintance over three days without even going halfway through my quota. On Sol xhigh. I love hearing about new models, but every time I just don’t know why I should use something worse. I tried some random model on fireworks a week ago, and it immediately went i…
Re: GLM-5.3-Flash
#544Earlier quoted context omitted.
I blocked Z.ai as soon as they were loading 10 different external providers including Alibaba who was just proven to execute silent sound fingerprinting mechanisms.
> silent sound fingerprinting mechanisms I hope you also block the other half of the internet because sure as hell google and facebook are way worse than that.
Re: GLM-5.3-Flash
#545Holy shit, is this model really that bad??? Just asked it a question via the custom opencode go endpoint routed over cloudflare ai gateway doesnt show me the correct models. My fault was that I set " https://opencode.ai/zen/go " as endpoint and tried my-gateway.com/custom-ocgo/v1/models, turns out I had to add /v1 to the opencode url and leave it on my -gateway.com. But first it told me that opencode is not on the co…
> I mean thats not even the level of LLama2 7b ... I think you answered your own question. Must be something wrong.
Re: GLM-5.3-Flash
#546Is anyone actually tried it in agentic coding (claude code loops)? Are apple silicon macs (M5 Max) capable of working with that model? what was the tps?
I've just used it for a fairly complex refactoring of the UI in a SwiftUI / AppKit app. It managed the refactoring in blazing colors, and the resulting UI looked really good. It was also quite fast. I'm impressed.
Re: GLM-5.3-Flash
#547Earlier quoted context omitted.
If the site can be viewed in the EU, it has to follow EU rules. Not different from other countries. What do you think why the normal polymarket site is blocked for US users.
I’ve got a bunch of sites that can be viewed from (checks notes) the internet. If people in Europe don’t like that, they can block it, or choose not to visit it. In the meantime, you’ve proven exactly what I first said, which is that they are claiming it applies worldwide. Here I am, just putting a site on the internet, and you’re telling me I have to follow EU laws concerning it. Nope. I don’t.
Ar eyou demanding that your rights are valid globally? That's more than the EU want with the GDPR.
BTW if you are concerned about the GDPR it means you track your visitors and store personal data about them otherwise you wouldn't have any problem.
Re: GLM-5.3-Flash
#548Good bicycle, good pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Re: GLM-5.3-Flash
#549Earlier quoted context omitted.
I've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side. The tradeoff is time (especially on RDMA4 hardwa…
I can't get 3.8 to exit thinking loops. It will just think and think and think on the most trivial topics. I wanted it to port a speed test powershell script to c#. Claude opus 5 completes it under 60 seconds. I let 3.8 churn about 6 different times for 30+ minutes and it never wrote a single line of code to disk. It wrote lots of lines in thinking. unsloth/Qwen3.8-27B-GGUF UD-Q3_K_XL DSH (pi) Any tips?
Try the same prompt with a larger quant (even if it runs very slowly because the model no longer fits in VRAM) & see if Qwen does better - if so, there’s your answer.
Re: GLM-5.3-Flash
#550On what hardware do they run this? If I'll visit Shenzen, can I buy these chips? I'd much rather at this stage give money to any-other-manufacturer-than-Nvidia. I have a use case where I have to run local models, can't share data offsite.
Its probably the huawei ascend 910, they are using.