Earlier quoted context omitted.
This is a problem I find with opus is will spend so long thinking then going “but wait what if” To point where I stop it and simple tell it to “start writing code you can work it out as you go along” Seems writers block also effects LLM
Fable was 20 times worse on that. It's clear it was the vibe coding model, as like no other model before, fully turned you into his assistant instead of the other way around.
GLM-5.2 is the new leading open weights model on Artificial Analysis
51–60 of 476 posts
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#52Earlier quoted context omitted.
I’m not that interested in models that I can’t run on my desktop for ~0€, which is my AI budget.
Electricity cost seems to be about $30/month for a 32B model on a GPU. It's probably better on Apple hardware. https://github.com/QuantiusBenignus/Zshelf/discussions/2 Not accounting for hardware, of course :)
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#53[0]: https://aibenchy.com/compare/deepseek-deepseek-v4-flash-high...
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#54Earlier quoted context omitted.
After having got a taste of Fable 5 for me Opus 4.8 doesn't cut it any more -- and I don't know how to put this, I don't know if it's just me, but it's rhetorical flourishes are starting to really grate on me, never mind that it is at times deliberately weasel-wordy and economical with the truth until pressed. Opus 4.8 is definitely a stronger coding agent than DeepSeek 4.0 or Kimi 2.7 succeeding where they flounder…
You are not alone. How about GPT 5.5? Does it come close to Fable 5?
Review the commits with both Claude and GPT 5.5 Xhigh. You can see that Fable is still sloppy(er) compared to GPT. You can test it the other way around as well(drive the dev with GPT and review with GPT and Claude). You get the same result Claude has an edge though and that’s on building more beautiful user interfaces.
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#55> On the Intelligence vs. Cost per Task Pareto Frontier: GLM-5.2 is on the Pareto frontier of the Intelligence vs Cost per Task chart, with the lowest cost per task among models at its intelligence level. GLM-5.2 costs ~$0.46 per task, compared to GLM-5.1 ($0.25), Kimi K2.6 ($0.31), MiniMax-M3 ($0.18) and DeepSeek V4 Pro (max, $0.05) am i missing something?
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#56QWEN 3.6 27b is already pretty good, but it should be possible to get a better option now that runs in the same hardware, right?
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#57Correct me if I'm wrong, but neither DeepSeek nor GLM have image input modality. This makes them less useful when looking at UIs, photos, screenshots, etc. doesn't it? Or do they have alternate ways of doing so?
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#58It’s expensive, and not as capable as the frontier models, but would have some pretty big benefits around privacy and agency.
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#59It seems to really be a nice step-up and is getting quite close to the frontier. I wish they'd start focusing on the reasoning efficiency now, though. I have a simple (relatively) test task to evaluate LLMs: writing a simple math evaluator library in Nim (it's about 400-600 lines total max), and GLM 5.2 (xhigh which maps to max effort) spent over 15 minutes (!) reasoning, spending about 45k tokens, before it finally…
If you want reasonable token usage, you need to run it GLM 5.2 at High. There is little drop in quality from Max to High (for most tasks). And it cuts token usage by 2 a 2.5x. GLM 5.2, Max is really something you only need for complex tasks.
In essence, GLM 5.2 is Opus 4.8 its little brother, at a way, WAY cheaper price.
There has been really no training on Opus models going on, really, none i tell you! /sarcasm
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#60Knowing very little about how to run these, how close are we to medium or larger businesses starting to buy hardware to run models like this to keep the models local? It’s expensive, and not as capable as the frontier models, but would have some pretty big benefits around privacy and agency.
Years.
Even Microsoft said they don't have enough for Github and need to call Amazon.
Getting a few even at decent prices is hard. Unless the shortages goes down...