Earlier quoted context omitted.
I believe it is unlikely. (Not because I do not believe NSA is hoarding 0-days, but for many other reasons.) I'm curious: to any professional vulnerability researchers reading this, what do you think?
[flagged]
GLM-5.3: Frontier coding with emergent cyber capabilities
591–600 of 626 posts
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#592Earlier quoted context omitted.
> Anyone who knows anything realises banning things is a) impossible and Maybe "It's really hard" is more accurate? We (humanity) for most part basically agreed to ban the usage of various chemical weapons in wartime, which seems to have drastically reduced the usage of it, even though it's still used by shit actors today from time to time. But it's hard to deny that usage didn't decrease after banning it, which make…
Chemical weapons are not used not because some agreements - they just messy and only good for killing civilians. Also contaminate area and might also kill your own personnel. If you look at all other banned weapons they are all used against Ukraine by Russia and nobody gives a damn.
We've literally destroyed countless of supplies of chemical weapons because of agreements about not using them, because we all agree they're absolutely horrible:
> The OPCW administers the terms of the CWC to 192 signatories, which represents 98% of the global population. As of June 2016, 66,368 of 72,525 metric tonnes, (92% of chemical weapon stockpiles), have been verified as destroyed.[40][41] The OPCW has conducted 6,327 inspections at 235 chemical weapon-related sites and 2,255 industrial sites. These inspections have affected the sovereign territory of 86 States Parties since April 1997. Worldwide, 4,732 industrial facilities are subject to inspection under provisions of the CWC.
CWC = Chemical Weapons Convention
> If you look at all other banned weapons they are all used against Ukraine by Russia and nobody gives a damn.
If only I had mentioned something about "shit actors using chemical weapons", basically predicting this exact response. But alas, here we are with zero defense about such a sharp point you made.
Of course even agreements are only agreements. What's valuable is what happens afterwards, but it'll take a while to get there in the context of war crimes by Russia and/or Ukraine, but it will be investigated once the conflict is over.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#593This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#594Earlier quoted context omitted.
The harness is the agent. LLM's can be asked to output things in JSON for example. The LLM then literally asks for things like "execute this cmd" or search/replace this string. The LLM outputs text, but in a deterministic format that can be parsed. The harness calls the LLM, exposes tools, executes tools the LLM asks for, gates tool use based on security controls. It's the runtime that the agent uses to do work.
Are the LLM and agent the same thing? Why different nouns ?
The agent/harness is the sotware that leverage this "dumb" autocompletion engine to do useful things by sending the good input to the model and doing useful things with the output.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#595Earlier quoted context omitted.
The public that buys the stock at ipo at this inflated valuation and unwittingly buys indexes which include it (as with spacex).
All of this is public info though, and pretty well publicized at that.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#596OpenAI and Anthropic need to just go ahead and give people access to the cyber models. Otherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.
They won't. People will notice that the models are overhyped once they can test them.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#597Missing multimodal again? It is so valuable in practise to be able to have the models see screenshots - I guess if they aren't in the benchmarks then nobody will focus on it. But it completely nixes these for some of my main use cases.
Probably not what you're after, but I've considered having a separate small mm-model act as a seeing-eye dog for the bigger more capable one.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#598Their coding plan switched to credits, didn’t it? What are the rate limits like, compared to Anthropic or Kimi K3? I remember trying their Coding Plan out before the change and the 5 hour limits felt too restrictive then even for light/medium work, especially cause of the whole peak and off-peak thing: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model... Nowadays, I’d probably go with their Max plan if the…
Currently I am on the new max plan with the 5hr limits and weekly limits, I can't speak for the credit plan. Using it exclusively in zcode because of the usage multiplier + the harness is genuinely good.
Easily do a billion tokens per week on my limits and generally have no problems with limits, however if I use it during peak hours I will hit the 5hr limit super fast even in zcode.
Zcode gives much more usage: Normal hours 1x -> 0.67x usage multiplier Peak hours 3x -> 2x usage multiplier
Peak lines up with the afternoon for me and I prefer coding morning/night so its not really a problem + I have the codex $20 plan and opencode go so I can always use other subs during peak hours.
If the peak hours are your main work hours (its a 4 hour window) then the value is going to be MUCH lower, especially in CC or other harnesses (1/3rd the usage limits is harsh).
It does also change a bit depending on demand, so I recommend getting the max plan because you get priority access if you really like the model, it would be a 10/10 recommendation for me if it had vision but rn its mainly useful for backend or throw away internal tools where IDK about the UI as much.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#599Earlier quoted context omitted.
To carry on this analogy - do test prep workbooks make you meaningfully more competent in general, or is it benchmaxing? (Versus studying textbooks for a similar time, of course.)
All machine learning is benchmaxxing. It's a big problem.