Live data from Hacker News

GLM-5.3-Flash

z.ai

291–300 of 605 posts

Re: GLM-5.3-Flash

#291
post #87

If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute). I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends…

We need to figure out what the real pricing is for a going concern. Right now, everyone is subsidizing and discounting to grow (or maintain) market share. The big question is whether the steady state, market derived inference pricing is above or below what we’re seeing today. I honestly don’t know. Anthropic had said that inference is profitable, but they’re clearly not yet profitable overall with training and buildo…

Has nobody from any of the companies hosting open weights models released detailed information on how much it really costs?

Re: GLM-5.3-Flash

#292

Earlier quoted context omitted.

Yeah, yeah. BUT, will the Sandero be ... load-bearing ? :)

It can bear the load of a few people, at least.

Hey, I can at least say you will get more value (or, at least more predictable value) from a Dacia than from Anthropic's tokens: "Upgrade to [SUPER DUPER] for 5x the [TOKEN-SERF PACKAGE] token use!"

Actual net work doable with/intelligence supplied by the [TOKEN-SERF PACKAGE]: Unknown. Fluctuating.-

Re: GLM-5.3-Flash

#294
post #170

Earlier quoted context omitted.

The cloud stuff is definitely a much better economic value, but I would argue: 1. You learn a lot more running this stuff yourself (especially since you can poke at its internals if you're interested or watch the reasoning chain.) Just being a consumer of this stuff doesn't really teach you much about it other than model & harness specific tricks that become obsolete pretty quickly. (IE, your Claude.md from 6 months…

3.privacy Any organisation or individuals not wanting to have their sensitive data flowing away (either because of trade secret or data protection laws)

Or good old fashioned privacy.

There’s no law or business advantage preventing me giving my financial transaction and medical info to Google/Anthropic/OpenAI but I just don’t want to.

Re: GLM-5.3-Flash

#295
post #200

Earlier quoted context omitted.

It's not like US companies don't do the same either.

It’s implied that they do, but don’t have the balls to tell you they do.

They tell you, and allow you to opt out in certain plans.

Re: GLM-5.3-Flash

#296

Earlier quoted context omitted.

It can bear the load of a few people, at least.

Hey, I can at least say you will get more value (or, at least more predictable value) from a Dacia than from Anthropic's tokens: "Upgrade to [SUPER DUPER] for 5x the [TOKEN-SERF PACKAGE] token use!" Actual net work doable with/intelligence supplied by the [TOKEN-SERF PACKAGE]: Unknown. Fluctuating.-

> you will get more value (or, at least more predictable value) from a Dacia

I picked up an older Dacia Sandero for cheap a few years back - it's the best money I've ever spent on a car, hands down. That car does not quit.

Re: GLM-5.3-Flash

#297

Earlier quoted context omitted.

The model weights are MIT licensed.

I'm talking about the Z.ai service specifically.

As others have mentioned, nothing's stopping any other major provider from offering it. Given its popularity, you can guess how that'll develop. So, overall, irrelevant.

Re: GLM-5.3-Flash

#298

Earlier quoted context omitted.

they have 2 interfaces each so you typically daisy chain them

That will hurt latency and latency is very important for good tensor-parallelism performance

I think I need one now that I have four - reduce is ring-oriented and still works I believe, but IIRC you only get 200gbps if you use _one_ of two connectx ports.

Re: GLM-5.3-Flash

#299
post #228

Earlier quoted context omitted.

I've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side. The tradeoff is time (especially on RDMA4 hardwa…

What harness? I've had similar results as _Implicated_ said above - it's not done well in any of the tests I've tried with it. I currently have it hung off DS4Flash as a pseudo-vision tool and subagent only because of this.

I'm using pi inside a self-made harness. I've found going super lightweight with context (AGENTS.md is maybe 20 lines) and letting the model discover what it needs to gives the best results.

Re: GLM-5.3-Flash

#300

So the vagueposting by googlers about Ox Alpha was just... what exactly? Like I get that they have to be careful about comms, but surely senior members of the team can clarify when something is NOT them, when everyone is gosspiing it is them.

Trolling. GLM is heavily distilled from Gemini.

Source? GLM is great for coding and Gemini is barely useful in coding, to be generous.

I highly suspect that the Gemini Google uses internally is very different from what they offer in Antigravity.

Post reply on HN