Why is GLM 5 more expensive than GLM 4.7 even when using sparse attention?
There is also a GLM 5-code model.
171–180 of 540 posts
Why is GLM 5 more expensive than GLM 4.7 even when using sparse attention?
There is also a GLM 5-code model.
It's looking like we'll have Chinese OSS to thank for being able to host our own intelligence, free from the whims of proprietary megacorps. I know it doesn't make financial sense to self-host given how cheap OSS inference APIs are now, but it's comforting not being beholden to anyone or requiring a persistent internet connection for on-premise intelligence. Didn't expect to go back to macOS but they're basically the…
Yeah that sounds great until it's running as an autonomous moltbot in a distributed network semi-offline with access to your entire digital life, and China sneaks in some hidden training so these agents turn into an army of sleeper agents.
very smart idea!
Earlier quoted context omitted.
> Didn't expect to go back to macOS but their basically the only feasible consumer option for running large models locally. I presume here you are referring to running on the device in your lap. How about a headless linux inference box in the closet / basement? Return of the home network!
Apple devices have high memory bandwidth necessary to run LLMs at reasonable rates. It’s possible to build a Linux box that does the same but you’ll be spending a lot more to get there. With Apple, a $500 Mac Mini has memory bandwidth that you just can’t get anywhere else for the price.
You want the M4 Max (or Ultra) in the Mac Studios to get the real stuff.
Here is the pricing per M tokens. https://docs.z.ai/guides/overview/pricing Why is GLM 5 more expensive than GLM 4.7 even when using sparse attention? There is also a GLM 5-code model.
I'd say that they're super confident about the GLM-5 release, since they're directly comparing it with Opus 4.5 and don't mention Sonnet 4.5 at all. I am still waiting if they'd launch GLM-5 Air series,which would run on consumer hardware.
Earlier quoted context omitted.
That's just whataboutism. Why shouldn't people talk about the various ideological stances embedded in different LLMs?
Why do we hear censorship concerns only when it comes Chinese models? Why don't we hear similar stances when Claude or OpenAI releases models? We either set the bar and judge both, or don't complain about censorship
why don't they publish at ARC-AGI ? too expensive?
Earlier quoted context omitted.
Not going to call $30/mo for a github copilot subscription "cheap". More like "extortionary".
Yeah it's funny how the needle has moved on this kind of thing. Two years ago people scoffed at buying a personal license for e.g. JetBrains IDEs which netted out to $120 USD or something a year; VS Code etc took off because they were "free" But now they're dumping monthly subs to OpenAI and Anthropic that work out to the same as their car insurance payments. It's not sustainable.
So whether you pay Claude or GitHub, Claude gets paid the same. So the consumer ends up footing a bill that has no reason to exist, and has no real competition because open source models can't run at the scale of an Opus or ChatGPT.
(not unless the EU decides it's time for a "European Open AI Initiative" where any EU citizen gets free access to an EU wide datacenter backed large scale system that AI companies can pay to be part of, instead of getting paid to connect to)