Whether or not distillation matters a small amount or a big amount, still interesting:
https://www.whitehouse.gov/presidential-actions/2026/06/nati...
791–800 of 965 posts
Whether or not distillation matters a small amount or a big amount, still interesting:
https://www.whitehouse.gov/presidential-actions/2026/06/nati...
Earlier quoted context omitted.
I'm talking about using mechanistic interpretability to see the model's intent. If it is deliberately using compromised libraries to weaken some code's security, there's going to be a signal in its hidden activations that it's doing so. Finding these kinds of activations is something Anthropic is actively researching [1] but they're the only ones who can use those techniques to see Claude's intent. On the other hand,…
Except that the model doesn't hold any malicious intent when doing it. No "deception", no "sabotage".
If it does have grounding, and can therefore see that it's introducing vulnerabilities to the code it's generating, yet does so anyway... I suppose we could invoke Hanlon's razor, but if the model is that incompetent, it probably isn't the right tool for the job regardless of its provenance.
That said, we aren't talking about incompetent models, we're talking about models sabotaging projects due to hidden motives. My point, again, is that those motives could potentially be revealed with open-weight models, in a way that will never be possible with closed models (barring some sort of legislation requiring independent third-party interpretability audits, which I suppose is in the realm of possibility).
The 2 things people need to remember: 1) China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China. 2) Ignoring the models containing false information, they are incredible. But you should be scared of running inference via the model creators directly. If you think your data is safe compared to running it via mod…
The vast majority of Westerners will interact with ChatGPT/Claude/Google in a browser. They'll use these models to try to save some money coding.
The post hits it spot on with unequal access to the models in terms of security. I'm developing OSS where security is important for the user ... but the frontier models like GPT 5.6 and Fable flake out and state that I cannot get the info/access. This is extremely lopsided I'll have to resort to GLM 5.2/K3 to ensure that those security issues (hopefully) are resolved properly. For OSS, this is one of the most counter…
Part of me wonders if the US Government is muzzling Anthropic and OpenAI so they can stockpile NOBUS exploits: https://en.wikipedia.org/wiki/NOBUS There would be a decently large incentive to restrict these models if they could be used to patch (or discover) dangerous payloads. In larger projects like Windows or Chrome, there might still be dozens of unpatched exploits that are too subtle to catch with smaller models…
Even during the pre-Snowden heyday of US cyber supremacy, these capabilities were barely part of the thought process of White House officials.
Earlier quoted context omitted.
Is that a good faith question? Please check the site guidelines.
uses rhetoric counterparty uses rhetoric in response "Hey! No fair!"
There's plenty of room for interesting discussion on what it takes to make green energy work. "What, are you against freedom?!" is not that.
Earlier quoted context omitted.
Part of me wonders if the US Government is muzzling Anthropic and OpenAI so they can stockpile NOBUS exploits: https://en.wikipedia.org/wiki/NOBUS There would be a decently large incentive to restrict these models if they could be used to patch (or discover) dangerous payloads. In larger projects like Windows or Chrome, there might still be dozens of unpatched exploits that are too subtle to catch with smaller models…
NOBUS exploits have rarely been a driving interest for elected officials. Trade restrictions and reciprocity are far more salient and legible. Most elected officials are only barely aware of what NOBUS exploits even mean. Even during the pre-Snowden heyday of US cyber supremacy, these capabilities were barely part of the thought process of White House officials.
I can believe that NOBUS and other backdoors were ignored for a long time, but I have a hard time believing that it's being ignored by the current administration.
The post hits it spot on with unequal access to the models in terms of security. I'm developing OSS where security is important for the user ... but the frontier models like GPT 5.6 and Fable flake out and state that I cannot get the info/access. This is extremely lopsided I'll have to resort to GLM 5.2/K3 to ensure that those security issues (hopefully) are resolved properly. For OSS, this is one of the most counter…
I'm afraid to trust technology coming from companies operating under the jurisdiction of a rogue aggressive nation that is continuously attacking other nations both (so called) ally and foe using economic and military actions.
Earlier quoted context omitted.
Outside of a few boarder disputes with India, I don't China has militarily attacked anyone since they got their ass handed to them by Vietnam (Sino-Vietnamese War 1979). So I think that rules out China.
Hong Kong would like a word...
Earlier quoted context omitted.
Are you describing the United States or China with this quote? It's hard to tell.
Outside of a few boarder disputes with India, I don't China has militarily attacked anyone since they got their ass handed to them by Vietnam (Sino-Vietnamese War 1979). So I think that rules out China.