So the vagueposting by googlers about Ox Alpha was just... what exactly? Like I get that they have to be careful about comms, but surely senior members of the team can clarify when something is NOT them, when everyone is gosspiing it is them.
GLM-5.3-Flash
281–290 of 605 posts
Re: GLM-5.3-Flash
#282Earlier quoted context omitted.
I don't understand how people don't consider this. Plus you're spec'd out of near-SOTA level in months. The only reasons to actually do this are a) you have a lot of dispensable income and are a hobbyist/tinkerer, b) you have real, legitimate privacy concerns or, relatedly, c) you're doing something you don't want to get flagged
Not everything is about pure cost. Maybe I don't want to sell my soul supporting the frontier labs because they are straight up pure evil?
Even if you trained your own model, you'd be committing some of the same sins, paying for the same hardware that drove it, etc. But if you're using some open model, you're standing on the shoulders of the same corrupt giants.
I feel like when people say this is due to moral reasons, it's to justify an expensive hobby.
Re: GLM-5.3-Flash
#283Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it. https://deepswe.datacurve.ai/ That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost…
I don't know how anyone can actually use Luna max on ANY real workload. I've had Sol orchestrate a bunch of Luna agents, these agents were explicitly given small chunks of larger objectives and they still filled their entire context windows with just reasoning tokens, until compaction hit, and then reasoning again. I've probably wasted a good 40% of my weekly usage on Luna Max agents just thinking and not writing a s…
Re: GLM-5.3-Flash
#284Earlier quoted context omitted.
To be fair, there is no 3 turns that I don't have to jump in into what Opus 5 is doing. There is either some regression or my prompting skills are so much worse now. Flash is not perfect and honestly some things depend on how big context do you keep. So I'm keeping like a really short context with my flash, but it works okay, even though it has a tendency to overthink, and yeah, I run it always in max effort mode.
Use Opus 4.8. 5 is absolute garbage. Don't use DS4 Flash in max effort mode. It's just spinning its wheels, in my experience (I have a harness for testing models with 25 real bugs/features/etc from my real projects that I measure outcomes against) DS4 flash does _worse_ with max effort. It will literally have the right approach and reason itself away from it.
Re: GLM-5.3-Flash
#285> with all of this traffic served on Chinese AI chips RIP Nivida shareholders
This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.- Further quote: "Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference effic…
Re: GLM-5.3-Flash
#286Earlier quoted context omitted.
> Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms. OpenAI revoked my Cyber verification, along with many others, asked to reverify (i.e. give my biometric information to Persona), had me do it 8 times, just to find out several days later that they silently implemented a nationality whitelist, and my nationality didn't make it (and no, it's not a sanctioned country). The…
> just to find out several days later that they silently implemented a nationality whitelist, and my nationality didn't make it (and no, it's not a sanctioned country) How did you discover this? I opened the Persona tab once, closed it and the tab never opened ever again. "Precheck failed". What countries are banned? I'm from Brazil. I went as far as initiating an LGPD (brazilian GDPR) process against them due to thi…
A week later, I tried testing the TAC flow on my SO's account, which had never had a TAC attempt before. Selecting Georgia in the Persona iframe now says, "We are unable to verify identities in this country."
But I know this isn't a Persona limitation, as I verified with Anthropic using Persona the same day.
So what I think happened was this: OpenAI silently implemented a country whitelist on their end and revoked TAC for affected individuals who already had it, calling it a "technical issue". They forgot to disable those countries in Persona, so everyone just got a cryptic error. Then they disabled them in Persona too.
Interaction with support was AI with human names, which essentially just repeats what you said. And it ended with:
> I’m unable to provide additional details about verification outcomes, and Support cannot manually override the result. At this time, Trusted Access for Cyber verification does not support retries or appeals.
I even provided them my credentials, and support AI was basically: lol wat we're here to check technical errors, your credentials are of no relevance.
Re: GLM-5.3-Flash
#287Earlier quoted context omitted.
I get all that. Then alternatives are: - Grok - where I absolutely have 0 trust in X.ai's interst in "pushing humanity forward". - OpenAI and Anthropic - which seem to try to be building the biggest moat they can by pushing to ban open models. And at the same time want to be an Arbiter of what level of intelligence I can use. - Google and Meta - I don't need to talk about the practices of these companies. Yes, the te…
All the American companies you mentioned still follow American law and regulation. Skirting that blatantly has big consequences. Chinese companies do not follow American laws and there are absolutely no consequences for violating it. Moreover, the average American is not even aware of exactly what the legal/judicial environment is like in China. If your code and data is stolen, you can't fly to China and demand justi…
No, they don’t. This is an absurd statement to make in 2026.
Re: GLM-5.3-Flash
#288Earlier quoted context omitted.
Isn't this practically every TOS though? Nearly every TOS I've ever read has a "We can ban you for any reason, or no reason, are under no obligation to disclose any reason." line somewhere in it. HN's for example > We reserve the right, at our sole discretion, to change or modify portions of these Terms of Use at any time. > You acknowledge that Y Combinator may establish general practices and limits concerning use o…
> Isn't this practically every TOS though? Not even close. Even OpenAI and Anthropic aren't bad enough that they claim literal ownership of your inputs and outputs. > HN's for example You're not paying to use HN. Getting banned here has essentially zero consequences. If Z.ai uses its absolute powers to ban you because you wrote a review about them or something, then you lose actual money. This is especially relevant…
At most I suspect the A.I. providers will just come up with yellow banners like Anthropic did where naughty smut writers get put in the time out corner.
Re: GLM-5.3-Flash
#289Earlier quoted context omitted.
Another self-inflicted own courtesy of US government policy. While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.
Long term it's irrelevant. The only relevant thing is that there's lots of money in chips that can do high performance inference. You see all kinds of competitor products in development or already on the market even here in the US where there are no such restrictions. Cerebras comes to mind. It's natural and expected that eventually Nvidia will either have to keep way ahead or competition will catch up with specializ…
Curiously, there is not a single real CUDA competitor anywhere in the world. We almost had one with OpenCL, but all of the American stakeholders abandoned it right before the crypto/AI takeoff. All of which means that Nvidia sets their own margins, exploiting American investors and taxpayers while letting China avoid their dominance. So the American economy subsumes the bulk of Nvidia's arbitrarily-priced debt, and the Chinese economy can direct SOEs to pour billions in liquid cash into real GPGPU research.
I'm an American and I'm pretty fond of Nvidia, but Jensen was right about this policy; it gives China everything they need to actually replace CUDA. It's reminiscent of America's attempts to deprive China of ARM and Texas Instruments IP, only to end up swimming in unlicensed clones after refusing to sign an IP deal.