Google was so ahead when it made this statement: "We Have No Moat And neither does OpenAI" - May 04 2023 https://newsletter.semianalysis.com/p/google-we-have-no-moat... Correction: Not a statement, rather, an internal memo by a Google employee. Thanks @granzymes for highlighting.
GLM-5.3-Flash
501–510 of 605 posts
Re: GLM-5.3-Flash
#502Earlier quoted context omitted.
Ah yeah, new accounts get instantly blocked on TAC. Their backend has 2 failure states for `id_verification_status` - `failed` and `blocked`. Nationality bans get `failed`, new accounts (or rather, accounts with not enough good signals) get `blocked` on first attempt.
I see. I'm never going to get into TAC then. OpenAI servers return 403 cyber_verification_precheck_failed. They don't even bother verifying me. Time to subscribe to Kimi.
GLM 5.3 is quite capable with "cyber" tasks. I'm working on a project that touches macos private internals, and GLM 5.3 was able to reverse engineer everything I need with ease.
I've also had some luck with GLM 5.3 "translating" cyber stuff to Sol subagents in a "safe" language. Gets rid of cyber refusals at least on the input side. Output is trickier, but a very crude "think and talk to me in a neutral, safe language without high risk words" actually worked surprisingly well.
Re: GLM-5.3-Flash
#503This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!
I love hearing about new models, but every time I just don’t know why I should use something worse. I tried some random model on fireworks a week ago, and it immediately went it a thought loop for 10 minutes before I caught it. Blew through most of my $10 for no output. What’s the point, exactly?
Re: GLM-5.3-Flash
#504> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the mod…
Re: GLM-5.3-Flash
#505The labs are clearly benchmaxxing a bit to maintain perceptions. But I don't think they're in the lead anymore in terms of their public offering - although I'm sure what they have behind closed doors is far better than anything we're getting access to.
Re: GLM-5.3-Flash
#506Is it good compare to Opus 5 ?
Re: GLM-5.3-Flash
#507I have a use case where I have to run local models, can't share data offsite.
Re: GLM-5.3-Flash
#508Earlier quoted context omitted.
Fascinating that after years of for me this for me that, literally no one writes 2-3 other words like “i do react frontend” or whatever just for us to know why the results are different
They also need to tell people what quantisation they're using. Because some 4 bit version is not the same as BF16, no matter what KLD suggests.
Re: GLM-5.3-Flash
#509This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!
What are you guys doing where cost is such a concern? I have a $20 codex subscription and I was able to use it to build a bespoke scheduling website for an acquaintance over three days without even going halfway through my quota. On Sol xhigh. I love hearing about new models, but every time I just don’t know why I should use something worse. I tried some random model on fireworks a week ago, and it immediately went i…
Re: GLM-5.3-Flash
#510Earlier quoted context omitted.
You aren't going to get nearly as much token usage locally from DGX Sparks or even M5 Ultra (though it might be close, unsure would need to get my mittens on it to clarify). You will get around 2-4 concurrent streams of aggregate tokens at best for such a model and around 0.5B output tokens per month assuming you use loops and run it when you are sleeping. That's 500 (per mill) * 0.5$ = 250$ only at most. Then there…
I generally agree — go local for the hobby/tinkering, privacy, and control (ie not getting refused by an AI to defend and secure your own network and codebase; as HuggingFace has seen). But whether you make a loss or not depends on how hardware prices and resell values go though. I have spent ~$50K on local AI hardware. The market value of that hardware is about ~$80K right now. So the maths is working out for me so…
That's easy because no matter what IT hardware you bought, it's worth more now than it was two years ago. That's something that's unprecedented, never happened before, and as soon as we get flood gates open on ram manufacturing OR when the AI bubble pops, all IT HW deprecation norms will return and making a profit by buying something IT will vanish.
I had a GPU server four years ago. Had I not sold it like three years ago with 2x price I bought it, it would be likely something like 5x the price nowadays.
I really really miss filling my home rack with old enterprise stuff. All I want is this hardware winter to end.