Earlier quoted context omitted.
> DeepSeek token prices are continuing to _increase_ One increase does not a trend make. And the current crop of models are now undercutting deepseek flash...
You can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right? Have you seen what's happening with Codex/Claude subscriptions? Deepseek raising API prices.. We've been getting subsidized tokens for some time now and as the hardware costs skyrocket these labs/people with inference compute are going to continue to clamp down.
GLM-5.3-Flash
311–320 of 605 posts
Re: GLM-5.3-Flash
#312> 320B total parameters and just 18B active parameters This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home. @edit: so many releases that I forgot to math. This fits j…
Speaking as someone who isn't really well versed in this, does 18B active parameters mean that you could potentially hold only the 18B parameters in RAM and stream the rest from a fast NVMe SSD for acceptable performance similar to how Colibri works? https://github.com/JustVugg/colibri
Generally, for local consumer use, these large MOE models are best for unified RAM systems like DGX Spark or Mac Studio.
Re: GLM-5.3-Flash
#313Re: GLM-5.3-Flash
#314You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…
>Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country.
[...] may cause harm to Anthropic, our users, or third parties, we reserve the right to remove or take down some or all of such Third-Party Content using, where appropriate, algorithmic and human review.
You may not export or provide access to the Services into any U.S. embargoed countries or to anyone on (i) the U.S. Treasury Department’s list of Specially Designated Nationals, (ii) any other restricted party lists identified by the Office of Foreign Asset Control, (iii) the U.S. Department of Commerce Denied Persons List or Entity List, or (iv) any other restricted party lists
>Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.
we will use Materials for model training when [...] your Materials are flagged for safety review to improve our ability to detect harmful content, enforce our policies, or advance our safety research.
>Vague prohibitions on discussing Z.ai, even my posting this comment violates it.
>Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan kiss your money goodbye.
To engage in any other conduct that restricts or inhibits any person from using or enjoying our Services, or that we reasonably consider exposes us—or any of our users, affiliates, or any other third party—to any liability, damages, or detriment of any type, including reputational harms.
Mind you, that's Anthropic's Terms of Use in Europe. I have zero doubts the TOS applied to the US is even worse and that merely mentioning your first born in a chat entitles them to a part of its soul.
Re: GLM-5.3-Flash
#315Earlier quoted context omitted.
I get all that. Then alternatives are: - Grok - where I absolutely have 0 trust in X.ai's interst in "pushing humanity forward". - OpenAI and Anthropic - which seem to try to be building the biggest moat they can by pushing to ban open models. And at the same time want to be an Arbiter of what level of intelligence I can use. - Google and Meta - I don't need to talk about the practices of these companies. Yes, the te…
> I don't believe that a future which OpenAI and Anthropic are pushing for has my best interest in mind. I don't believe in that either, but these totalitarian terms are absolutely unacceptable.
Re: GLM-5.3-Flash
#316Good bicycle, good pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
I only see a mostly blank page with a "Paste" button, a "URL" button, and and a "Preview" label.
Re: GLM-5.3-Flash
#317Re: GLM-5.3-Flash
#318Earlier quoted context omitted.
It’s implied that they do, but don’t have the balls to tell you they do.
They tell you, and allow you to opt out in certain plans.
And no, they won't tell you what their automated reviews consider "dangerous".
Re: GLM-5.3-Flash
#319Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…
Re: GLM-5.3-Flash
#320Earlier quoted context omitted.
You can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right? Have you seen what's happening with Codex/Claude subscriptions? Deepseek raising API prices.. We've been getting subsidized tokens for some time now and as the hardware costs skyrocket these labs/people with inference compute are going to continue to clamp down.
$40,000 GPU is like few pennies in sand. Only mildly hyperbolic. But a GPU fresh out of fab is $2000 after ASML, TSMC and inputs get their 50-75% margin, then somehow $40k laundered through US financialization / Nvidia margins. Commoditized GPUs shouldn't cost more than 1-2% current price once there's competition.