{
"providers": {
"opencode-go": {
"models": [
{
"id": "glm-5.3-flash",
"name": "GLM-5.3 Flash",
"api": "openai-completions",
"baseUrl": "https://opencode.ai/zen/go/v1",
"reasoning": true,
"input": ["text", "image"],
"cost": {
"input": 0.15,
"output": 0.5,
"cacheRead": 0.03,
"cacheWrite": 0
},
"compat": {
"supportsStore": false,
"supportsDeveloperRole": false,
"maxTokensField": "max_tokens"
},
"contextWindow": 1000000,
"maxTokens": 131072,
"thinkingLevelMap": {
"off": null,
"minimal": null,
"low": "low",
"medium": null,
"high": "high",
"xhigh": null,
"max": "max"
}
}
]
}
}
}GLM-5.3-Flash
351–360 of 605 posts
Re: GLM-5.3-Flash
#352Earlier quoted context omitted.
Source? GLM is great for coding and Gemini is barely useful in coding, to be generous. I highly suspect that the Gemini Google uses internally is very different from what they offer in Antigravity.
No, it's the same internally and externally. Gemini 3.7 Flash is a pretty great model IMO. You shouldn't compare it to Opus, Sol, K3, etc since it's a much smaller model but it's a little better compared to Sonnet, Luna or Terra, etc.
Re: GLM-5.3-Flash
#353Good bicycle, good pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
With GLM 5.1 and 5.2, the big problem was tool calling and long-horizon coherency. 5.3 was more trustworthy at the cost of longer thinking traces, and now Flash seems to improve on it once again with a more concise, smaller model. As long as there aren't any noticeable regressions, I could see myself defaulting to this for >90% of my day-to-day coding work.
Re: GLM-5.3-Flash
#354Re: GLM-5.3-Flash
#355Re: GLM-5.3-Flash
#356Earlier quoted context omitted.
> just to find out several days later that they silently implemented a nationality whitelist, and my nationality didn't make it (and no, it's not a sanctioned country) How did you discover this? I opened the Persona tab once, closed it and the tab never opened ever again. "Precheck failed". What countries are banned? I'm from Brazil. I went as far as initiating an LGPD (brazilian GDPR) process against them due to thi…
When they initially revoked TAC for a bunch of users due to a "technical error", the TAC verification flow opened a Persona iframe where you select the document country first. I was able to select it there (Georgia, in my case), went through the entire flow, and then got locked out after 8 attempts. Persona itself was successful end to end, so it failed somewhere on the OAI side. Other users then started reporting th…
For the record, I just made a second account, got verified by Persona and still didn't get into TAC. No "we are unable to verify identities in this country" message. No mention of my country whatsoever. Persona verification was successful.
What else do they want from us?
Were it not for Z.ai's obnoxious terms, I would have switched to them already...
Re: GLM-5.3-Flash
#357Earlier quoted context omitted.
This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.- Further quote: "Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference effic…
Anyone knows what are those Chinese chips? Can they be bought? (Assuming im not i the US, And actually im in a 3rd world country).
https://en.wikipedia.org/wiki/HiSilicon#Ascend_910
https://medium.com/@huaweiclouddevelper/a-brief-introduction...
Re: GLM-5.3-Flash
#358Earlier quoted context omitted.
> I don't believe that a future which OpenAI and Anthropic are pushing for has my best interest in mind. I don't believe in that either, but these totalitarian terms are absolutely unacceptable.
If only you could grab those models and host them literally anywhere else where you wouldn't be subject to those terms. Damn. Maybe we'll have to wait for someone to invent something like open download of model weights.
Re: GLM-5.3-Flash
#359At the current 50%-off GLM-5.3-Flash price ($0.075/M input, $0.25/M output; cached input $0.015/M), surprisingly, roughly $400–900/month would buy token throughput comparable to fully exhausting Claude Max 20×