This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…
OpenAI and Anthropic are both seeking trillion IPOs, while Chinese labs are pumping out open-weight models that are free for US providers to host and monetize. These Chinese models cost less of US SOTA models to run, even if they are less capable. Providers can just run them, offer cheap tokens, and pocket the margin. I just don't see how you justify a trillion valuation for US AI labs when the underlying models are…
GLM-5.3: Frontier coding with emergent cyber capabilities
101–110 of 626 posts
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#102No Hugging Face link yet. I wish they would release it under a true FOSS license. Kimi and QWEN are now moving on to a restricted-usage license, which, although is still better than the proprietary American models, is a step back from the open source Chinese LLM culture.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#103Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#104I might be just reading my positive bias into that text, but is it possible that it is written less like SV marketing hype trash and more like researchers wrote it? It does feel like it respects both me and my time. Thank you, Z.AI. Amazing what difference it makes when the top of your org are actual university professors.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#105Missing multimodal again? It is so valuable in practise to be able to have the models see screenshots - I guess if they aren't in the benchmarks then nobody will focus on it. But it completely nixes these for some of my main use cases.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#106Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#107> Scaling post-training is all we did for GLM-5.3. Love this opening line. And wow, great results. > As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.
does this suggest 5.3 is the same # of parameters as 5.2?
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#108Earlier quoted context omitted.
Anthropic needs to teach Opus how to speak English again, because Opus 5 seems to have forgotten. Utterly incoherent a lot of the time. They seem to be so busy scare-mongering and cooking up guardrails and watermarks that they haven't noticed that their models are getting weird.
It still knows how to speak English. When I tell it to explain something in plain language, it generally does a very good job. The weird thing is that those instructions don't persist: it lapses back into Claude-speak pretty much every turn no matter how hard I try to instruct it not to. (In my case "it"=Fable; I assume Opus is similar.)
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#109Earlier quoted context omitted.
We applied for the cybersecurity approval via the form and got approval back in less than an hour. Have you… tried?
You must be a 5000 person company with an existing enterprise contract to get approved that fast. That sounds like a 15 minute SLA agreement. Individuals no matter how qualified about cybersecurity, are ghosted
However, even being in the cybersecurity programme, Fable refuses to answer prompts that it determines could be even tangentially related to cybersecurity. In fact, for a while, I was unable to use Fable with any prompt, as it recalled from memory that I was a cybersecurity professional, which triggered the refusal even for simple prompts like asking for a chili recipe.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#110Apple will release M7 MacBook Pros / Mac Minis next year, and they will be able to run free LLMs locally at native speed. All software developer notebooks will be replaced to run local models, saving a lot by cancelling Claude Code subscriptions. Developers win. Apple stocks will be rocketing. Everything else will go down. You're welcome.