If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute). I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends…
We need to figure out what the real pricing is for a going concern. Right now, everyone is subsidizing and discounting to grow (or maintain) market share. The big question is whether the steady state, market derived inference pricing is above or below what we’re seeing today. I honestly don’t know. Anthropic had said that inference is profitable, but they’re clearly not yet profitable overall with training and buildo…
GLM-5.3-Flash
291–300 of 605 posts
Re: GLM-5.3-Flash
#292Earlier quoted context omitted.
Yeah, yeah. BUT, will the Sandero be ... load-bearing ? :)
It can bear the load of a few people, at least.
Actual net work doable with/intelligence supplied by the [TOKEN-SERF PACKAGE]: Unknown. Fluctuating.-
Re: GLM-5.3-Flash
#293How much is the “discounted” pricing they mention?
Re: GLM-5.3-Flash
#294Earlier quoted context omitted.
The cloud stuff is definitely a much better economic value, but I would argue: 1. You learn a lot more running this stuff yourself (especially since you can poke at its internals if you're interested or watch the reasoning chain.) Just being a consumer of this stuff doesn't really teach you much about it other than model & harness specific tricks that become obsolete pretty quickly. (IE, your Claude.md from 6 months…
3.privacy Any organisation or individuals not wanting to have their sensitive data flowing away (either because of trade secret or data protection laws)
There’s no law or business advantage preventing me giving my financial transaction and medical info to Google/Anthropic/OpenAI but I just don’t want to.
Re: GLM-5.3-Flash
#295Re: GLM-5.3-Flash
#296Earlier quoted context omitted.
It can bear the load of a few people, at least.
Hey, I can at least say you will get more value (or, at least more predictable value) from a Dacia than from Anthropic's tokens: "Upgrade to [SUPER DUPER] for 5x the [TOKEN-SERF PACKAGE] token use!" Actual net work doable with/intelligence supplied by the [TOKEN-SERF PACKAGE]: Unknown. Fluctuating.-
I picked up an older Dacia Sandero for cheap a few years back - it's the best money I've ever spent on a car, hands down. That car does not quit.
Re: GLM-5.3-Flash
#297Earlier quoted context omitted.
The model weights are MIT licensed.
I'm talking about the Z.ai service specifically.
Re: GLM-5.3-Flash
#298Earlier quoted context omitted.
they have 2 interfaces each so you typically daisy chain them
That will hurt latency and latency is very important for good tensor-parallelism performance
Re: GLM-5.3-Flash
#299Earlier quoted context omitted.
I've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side. The tradeoff is time (especially on RDMA4 hardwa…
What harness? I've had similar results as _Implicated_ said above - it's not done well in any of the tests I've tried with it. I currently have it hung off DS4Flash as a pseudo-vision tool and subagent only because of this.
Re: GLM-5.3-Flash
#300So the vagueposting by googlers about Ox Alpha was just... what exactly? Like I get that they have to be careful about comms, but surely senior members of the team can clarify when something is NOT them, when everyone is gosspiing it is them.
Trolling. GLM is heavily distilled from Gemini.
I highly suspect that the Gemini Google uses internally is very different from what they offer in Antigravity.