Earlier quoted context omitted.
I don’t understand why there isn’t public dataset for reasoning that can be improved by humans/llms like Wikipedia (ie with auto judging contributions etc).
For reasoning a manually-curated dataset is too small; you need to be able to automatically generate vast volumes of synthetic reasoning data with provably correct answers. That's presumably why Claude and GPT are so good at using Lean (the theorem prover), because they get fed a bunch of synthetic, verifiably correct training data.
GLM-5.2 is the new leading open weights model on Artificial Analysis
451–460 of 476 posts
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#452Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#453Earlier quoted context omitted.
It’s been amazing to see the arc of tech people going from “evil Disney, copyright is an abomination, information wants to be free” to “OMG copyright is inviolable and AI is taking money out of Plato’s descendants’ pockets!”
> taking money out of Plato’s descendants’ pockets Yeah, remind me - is it Plato's descendants that people are concerned about here, or is it every single author who had any work in Anna's Archive, any work published online, any work published on github, etc? I think that people are probably upset about the harm to living people who had their work stolen by Meta and other LLM companies - regardless of license, terms…
I’m not even disagreeing. I’m just saying the shift in attitude about copyright in the tech space has been sudden, dramatic, and really funny. Remember “you wouldn’t steal a car”? Today’s anti-AI tech contingent are enthusiastically embracing that false equivalence that we all laughed at 20 years ago.
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#454Earlier quoted context omitted.
Your reply doesn't seem to be in good faith. Please provide your formula for calculating effective per token cost. I am not sure why the small team argument is relevant. This is a crowded market, there are dozens if hundreds of third party inference providers in the world right now. I'm glad that's a good excuse that works on you but I'm not sure why the average user should care.
The formula is very easy. Go to the website of neuralwatt, and read ... 5$ = 1Kwh in power for non-subscription usage. For subscription usage you get ~50% more. Then you actually use the service and see how much tokens you use on average. You calculate the token use vs what you pay. And this gives you a stable number to compare different services and model with, if you want the token cost. This is basic school level…
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#455I have a question, as it happens: Do you think the benchmarks and models were trained on benchmark datasets to skew the results, even though in real-world applications we realize they're not that great?
Recent incident with the Rio 3.5 model clearly shows that many coding models are specifically trained/fine tuned for the benchmarks.
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#456Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#457Earlier quoted context omitted.
> taking money out of Plato’s descendants’ pockets Yeah, remind me - is it Plato's descendants that people are concerned about here, or is it every single author who had any work in Anna's Archive, any work published online, any work published on github, etc? I think that people are probably upset about the harm to living people who had their work stolen by Meta and other LLM companies - regardless of license, terms…
Sure, that’s the motte / bailey. Easy to point to living, starving writers who suffer grevious harm, in defense of perpetual copyright. Disney and others use literally this exact argument year after year. I’m not even disagreeing. I’m just saying the shift in attitude about copyright in the tech space has been sudden, dramatic, and really funny. Remember “you wouldn’t steal a car”? Today’s anti-AI tech contingent are…
If like, Disney did a 180 overnight and bought rights from Google to scan every writer's saved work in Docs with some flimsy legal argument then a person saying "wait doesn't copyright actually protect that" would make sense. Even if you were previously upset about them suing schools for using 80 year art.
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#458It seems to really be a nice step-up and is getting quite close to the frontier. I wish they'd start focusing on the reasoning efficiency now, though. I have a simple (relatively) test task to evaluate LLMs: writing a simple math evaluator library in Nim (it's about 400-600 lines total max), and GLM 5.2 (xhigh which maps to max effort) spent over 15 minutes (!) reasoning, spending about 45k tokens, before it finally…
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#459Earlier quoted context omitted.
The formula is very easy. Go to the website of neuralwatt, and read ... 5$ = 1Kwh in power for non-subscription usage. For subscription usage you get ~50% more. Then you actually use the service and see how much tokens you use on average. You calculate the token use vs what you pay. And this gives you a stable number to compare different services and model with, if you want the token cost. This is basic school level…
The irony of questioning someone's communication skills immediately after this exchange is hard to miss.
Re: GLM-5.2 is the new leading open weights model on Artificial Analysis
#460Earlier quoted context omitted.
> GLM 5.2 Max = Opus 4.8 Max in thinking behavior This is insane! I can't wait until technology progresses to the point we can run these things on consumer hardware!
Are there any indications that this will be possible? Consumer hardware will continue getting better but I can't see 512GB RAM in a MacBook Pro any time soon. I'm hoping linear attention techniques plus MoE will make breakthroughs in size/compression and throughput.
Could totally see this being a comment from a forum in like 1994 but swap out GB for MB and MacBook Pro to whatever the popular consumer pc was at the time