Earlier quoted context omitted.
Why not? Mythos level really doesn't seem that scary. And it would be a great way to take away the American labs international market. I think it would make strategic sense for them to release more capable models than what American labs are allowed to make available to the world. It would help them grow their global soft-power and be a destabilizing effect on the American economy.
It is fairly obvious to me that the open models are a form of "dumping" as far as the economics and the desired outcome from China's perspective. They get to watch as the US pours tons of money and talent into an industry, then prevent that investment from having any return. In 5 years we'll be on equal footing, China will have spent 1/1000th the money, and the only downside will be that they spent 5 years being 6 mo…
GLM 5.2 beats Claude in our benchmarks
411–420 of 559 posts
Re: GLM 5.2 beats Claude in our benchmarks
#412Re: GLM 5.2 beats Claude in our benchmarks
#413Earlier quoted context omitted.
> most halfway decent models can write damn good code for a fraction of the price. The difference is how the model is used. With Opus you can give it a long-horizon task (eg build an entire feature) and it will plan it out and implement it and almost always stay on task. This is what people mean when they say "agentic tasks" With the lessor models the code is fine, but they need something else to plan what needs to b…
I would say 3.5 flash is great if you use a good open harness. I use omp for that. The thing with Google is that they announce they have a great model, and that they have been testing it internally for half a year. I guess they don't care too much about who or how he uses it. I am still struggling how to deal with sub agents and different roles for each model. I still think Claude or Codex are overall better models,…
As for Fable: I used it as much as I could while we had it.
It was a step change over Opus with my work.
Re: GLM 5.2 beats Claude in our benchmarks
#414How to reconcile that with the recent, highly upvoted, article titled: "The gap between open weights LLMs and closed source LLMs" ? What explains it? Is TFA lying? Is the most upvoted comment here lying?
The article itself doesn't say "it's better", basically just says "in this one specific benchmark it beat Claude with Claude code". Mind you with multimodality it Opus still beat GLM 5.2 very handily in that same benchmark.
I can't find any contradiction and I don't see anyone lying directly. At most they lead you to imply false things, but they're not untrue at a literal reading.
Re: GLM 5.2 beats Claude in our benchmarks
#415Any good resources about this (also for setup and recommend config)?
Re: GLM 5.2 beats Claude in our benchmarks
#416Earlier quoted context omitted.
Because car loans can’t be used to buy computers
Surprising that the banking industry has not come up yet with the AI native consumer product loan for GPUs.
Re: GLM 5.2 beats Claude in our benchmarks
#417Earlier quoted context omitted.
There is no money made from these people though .. people who are using ChatGPT to plan for their next week-end or their next vacation aren't paying a $100 or $200 monthly subscription. As for non coder office workers (accountants, PMs, etc.), they use Microsoft or Google products which all integrate AI to some extent within their products - with RAG for Sharepoint to some basic AIs to generate text or automate work…
You’re making a good point. I don’t disagree with what you’re saying. But I think my point got lost. I don’t agree with “Software development is where money is made for these labs”. Coders will inevitably eat up the most tokens & buy the bigger $200 subscriptions because we want to keep working. But us coders are still the small minority of users. They aren’t counting on us to get to trillion dollar evaluations. They…
I know Google gives me free Gemini AI from my Google Drive plan. Microsoft probably already does too, didn't test. Apple is probably crafting some arrangements if not offering already.
My point is most people wont pay for AI. It will be bundled.
And I think AI is going to be free for all, with ads.
Re: GLM 5.2 beats Claude in our benchmarks
#418Re: GLM 5.2 beats Claude in our benchmarks
#419Earlier quoted context omitted.
Yeah, the funniest thing about everyone freaking out about Fable's capabilities recently was that for most of the stuff they were amazed by, you could get roughly the same result from DeepSeek Flash. I used to be obsessed with what's the best model. Then a while back when the new best model came out, I tested it on a task. I also tested its little brother (much smaller model from same company). They both completed th…
"Best model" discourse always remember me of my days in Monster Hunter with people who refused to consider playing with anything other than the meta set for their weapon and then proceed to immediately cart right at the beginning of the hunt :) With the wealth of models available (open source vs closed, api vs local), I find optimizing the cost-efficiency of your token consumption an important part of business-orient…