Earlier quoted context omitted.
Can you recommend any US based cloud providers?
In HuggingChat ( https://huggingface.co/chat ) you can test open models for free and even test specific providers. From there I collected the following US providers currently serving GLM 5.2: - Together ( https://www.together.ai/models ) - Fireworks ( https://fireworks.ai/models ) - Featherless ( https://featherless.ai/models )
GLM 5.2 vs. Opus
261–270 of 367 posts
Re: GLM 5.2 vs. Opus
#262I seriously dont' know all this big hullabaloo about one shot prompting. by definition, a single prompt wont' constitute the complexity of a software project. ergo, what you'll get is a series of assumptions made by the model based on preexisting code in its training corpus. I'd rather see a coding agent that can follow steps in a plan file to a T while following guardrails and adhering to the proper coding conventio…
- Vibes are too subjective, I want an actual A/B test!
- An A/B test is too limited, I want a benchmark! (You are here.)
- Those benchmarks never seem to be reliable, I just go on vibes.
Re: GLM 5.2 vs. Opus
#263Earlier quoted context omitted.
it wont happen, its all a money grab.
I think that LLMs will stay, but I also think we've plateaued and that big companies will fail and fall and we will have another years long "halt" of any real advancements coming to the public. Similar to how ML was all the hype about 12 years ago and then it submerged again for a couple of years.
One can hope. Probably an unpopular take here but I'm tired boss.
The software world has a huge backlog of things that can all be done with the tech we currently have, no breakthrough advancements needed, but none of it will get prioritized when we're all forced to run on the new and shiny treadmill. Ever since LLM hype its like the javascript culture of a new framework every 10 minutes has infected every other vertical of software development and I'm exhausted.
Re: GLM 5.2 vs. Opus
#264This implies Opus was potentially much (?) better value.
GLM cost a quarter but Opus was twice as fast. So we are already at GLM actually costing half when you compare on time, without even considering the extra effort and time it would take to get Opus-par results.
It's good to have cheaper options and very impressive to see the Chinese continue to set open standards in this field, but the article is maybe a little over-generous.
Re: GLM 5.2 vs. Opus
#265Earlier quoted context omitted.
The streetlight effect: > A policeman sees a drunk man searching for something under a streetlight and asks what the drunk has lost. He says he lost his keys and they both look under the streetlight together. After a few minutes the policeman asks if he is sure he lost them here, and the drunk replies, no, and that he lost them in the park. The policeman asks why he is searching here, and the drunk replies, "this is…
Sure, for casual evaluation, I agree. But are there serious analyses that are evaluating this kind of thing? I mean, these are the kinds of things I evaluate in my own work when a new model comes out, or when I'm evaluating a harness. But this is all very ad hoc and intuitional. I'd love to start bringing rigor to it, but I haven't found much prior art on this. In another thread someone said that's because it's proba…
Re: GLM 5.2 vs. Opus
#266Re: GLM 5.2 vs. Opus
#267Earlier quoted context omitted.
Taking a view from outside the USA, European companies just had Fable taken away due to US export controls, and before that Anthropic announced it is holding their data for 30 days. There is immediate value to these firms to build their infrastructure around an AI that won’t be pulled away from them. And outside of Europe, other countries are more price sensitive and don’t have the same fear of building relationships…
And you have that guarantee from Xi?
Re: GLM 5.2 vs. Opus
#268Earlier quoted context omitted.
Why not? Given a proper spec, you should absolutely be able to one-shot Excel, particularly if we put it at the level of complexity of, say, Excel 1.0 for Mac. Current models aren't capable of that, but that doesn't mean it's not possible.
The issue is not the models, the issue is that this method ws tried before, and humans suck at writing what they want. Developing in small increments allowing feedback was an answer to this issue. If you made models able to code to long spec, you would be left with the hard issue of having to write them.
Like if you show the LLM a page, can the LLM review the page and then spit out a review that is close to what a human would say about the page?
Re: GLM 5.2 vs. Opus
#269Re: GLM 5.2 vs. Opus
#270there is no comparison between glm 5.2 and opus. First for this glm 5.2 you need a big big resource and that big also came from money so instead you buy the opus subscription and enjoy.
you go to OpenRouter and pay
$0.98 / $3.08per 1M for GLM 5.2 vs $5 / $25per 1M for Opus.
GLM 5.2 gives you OPTIONALITY so you can run it locally, but you can still just pay somebody for it.