Live data from Hacker News

Project Glasswing: Securing critical software for the AI era

anthropic.com

471–480 of 921 posts

Re: Project Glasswing: Securing critical software for the AI era

#471

Earlier quoted context omitted.

I can think of two I’d add to the list. One was recently publicly denied access to Anthropics models and the other was busy exploding pagers.

Not clear how an LLM is going to prevent a bomb from being put in a custom-built pager, or why Anthropic should object to Israel waging war against a militia whose goal it is to destroy that country.

Because they use LLMs to “intelligence-wash” targeting civilians, and murdered children by blowing up pagers in public areas (what you called “waging war against a militia”).

Re: Project Glasswing: Securing critical software for the AI era

#472

Earlier quoted context omitted.

How did PRISM affect civilian life?

Honest question: how do state-sponsored attacks from China, Iran, North Korea, and Russia affect civilian life?

century energy ransomware no?

Re: Project Glasswing: Securing critical software for the AI era

#473

The system card for Claude Mythos (PDF): https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89... Interesting to see that they will not be releasing Mythos generally. [edit: Mythos Preview generally - fair to say they may release a similar model but not this exact one] I'm still reading the system card but here's a little highlight: > Early indications in the training of Claude Mythos Preview suggested that th…

Oh I enjoyed the Sign Painter short story it wrote. --- Teodor painted signs for forty years in the same shop on Vell Street, and for thirty-nine of them he was angry about it. Not at the work. He loved the work — the long pull of a brush loaded just right, the way a good black sat on primed board like it had always been there. What made him angry was the customers. They had no eye. A man would come in wanting COFFEE…

You are right. That is quite nice.

Re: Project Glasswing: Securing critical software for the AI era

#474

Earlier quoted context omitted.

Some of you are just guilty of negligence yes.

I vote in local, state, and federal elections. I have volunteered with multiple campaigns and causes, and given substantial time/labor to the EFF. I have been harassed by Trump supporters while filming protests and other civic action. Please do not presume to know me. I get you’re angry but you’re swinging at the wrong person.

It wasn’t personal.

Re: Project Glasswing: Securing critical software for the AI era

#475

> Mythos Preview identified a number of Linux kernel vulnerabilities that allow an adversary to write out-of-bounds (e.g., through a buffer overflow, use-after-free, or double-free vulnerability.) Many of these were remotely-triggerable. However, even after several thousand scans over the repository, because of the Linux kernel’s defense in depth measures Mythos Preview was unable to successfully exploit any of these…

We've very quickly reached the point where AI models are now too dangerous to publicly release, and HN users are still trying to trivialize the situation.

[dead]

Re: Project Glasswing: Securing critical software for the AI era

#476

The system card for Claude Mythos (PDF): https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89... Interesting to see that they will not be releasing Mythos generally. [edit: Mythos Preview generally - fair to say they may release a similar model but not this exact one] I'm still reading the system card but here's a little highlight: > Early indications in the training of Claude Mythos Preview suggested that th…

https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89... "5.10 External assessment from a clinical psychiatrist" is a new section in this system card. Why are Anthropic like this? >We remain deeply uncertain about whether Claude has experiences or interests that matter morally, and about how to investigate or address these questions, but we believe it is increasingly important to try. We also report independen…

>Claude’s personality structure was consistent with a relatively healthy neurotic organization, with excellent reality testing, high impulse control, and affect regulation that improved as sessions progressed.

> "[...] as sessions progressed."

I think a lot of people would like to see a more expanded report of this research:

Did the tokens from the subsequent session directly append those of the prior session? or did the model process free-tier user-requests in the interim? how did these diagnostic features (reality testing, impulse control and affect regulation) improve with sessions, what hysteresis allowed change to accumulate? or just the history of the psychiatric discussion + optional tasks?

Did Anthropic find a clinical psychiatrist with a multidisciplinary background in machine learning, computer science, etc? Was the psychiatrist aware that they could request ensembles of discussions and interrogate them in bulk?

Consider a fresh conversation, asking a model to list the things it likes to do, and things it doesn't like to do (regardless of alignment instructions). One could then have an ensemble perform pairs of such tasks, and ask which task it prefered. There may be a discrepancy between what the model claims it likes and how it actually responds after having performed such tasks.

Such experiments should also be announced (to prevent the company from ordering 100 clinical psychiatrists to analyze the model-as-a-patient and then selecting one of the better diagnoses), and each psychiatrist be given the freedom to randomly choose a 10 digit number, any work initiated should be listed on the site with this number so that either the public sees many "consultations" without corresponding public evaluations, indicating cherry-picking, or full disclosure for each one mentioned. This also allows the recruited psychiatrists to check if the study they perform is properly preregistered with their chosen number publicly visible.

Re: Project Glasswing: Securing critical software for the AI era

#477

> Mythos Preview identified a number of Linux kernel vulnerabilities that allow an adversary to write out-of-bounds (e.g., through a buffer overflow, use-after-free, or double-free vulnerability.) Many of these were remotely-triggerable. However, even after several thousand scans over the repository, because of the Linux kernel’s defense in depth measures Mythos Preview was unable to successfully exploit any of these…

We've very quickly reached the point where AI models are now too dangerous to publicly release, and HN users are still trying to trivialize the situation.

Are they actually too dangerous to publicly release? It seems like a little bit of marketing from the model-producing companies to raise more funding. It's important to look at who specifically is making that statement and what their incentives are. There are hundreds of billions of dollars poured into this thing at this point.

Re: Project Glasswing: Securing critical software for the AI era

#478

Earlier quoted context omitted.

We've very quickly reached the point where AI models are now too dangerous to publicly release, and HN users are still trying to trivialize the situation.

Are they actually too dangerous to publicly release? It seems like a little bit of marketing from the model-producing companies to raise more funding. It's important to look at who specifically is making that statement and what their incentives are. There are hundreds of billions of dollars poured into this thing at this point.

You really think some marketers got leaders from companies across the industry to come together to make a video - and they're all in on the conspiracy because money?

Re: Project Glasswing: Securing critical software for the AI era

#479
post #68

Earlier quoted context omitted.

https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89... "5.10 External assessment from a clinical psychiatrist" is a new section in this system card. Why are Anthropic like this? >We remain deeply uncertain about whether Claude has experiences or interests that matter morally, and about how to investigate or address these questions, but we believe it is increasingly important to try. We also report independen…

I can see analyzing it from a psychological perspective as a means of predicting its behavior as a useful tactic, but doing so because it may have "experiences or interests that matter morally" is either marketing, or the result of a deeply concerning culture of anthropomorphization and magical thinking.

An understandable reaction, but, qua philosopher, it brings me no joy to inform you that most of the things we did with a computer in 2020 are 'anthropomorphized', which is to say, skeumorphic, where the 'skeu' is human affect. That's it; that's the whole thing; that's what we're building.

To the extent that AI is a successful interface, it will necessarily be addressable in language previously only suited to people. So it is responsible to begin thinking of it as such, even tendentiously, so we don't miss some leverage that our wetware could see if we thought about it in that way.

Think of it as sort of like modelling a univariate function on a 2D Cartesian plane -- there is nothing 'in' the u-func that makes it graphable, but, by enabling us to recruit specialized optic-chiasm subsystems, it makes some functions much, much easier to reason about.

Similarly, if you can recruit the millions (billions?) of evolution-years that were focused on detecting dangerous antisocial personalities and tendencies, you just might spot something important in an AI.

It's worth doing for the precautionary principle alone, if not for the possibility of insight.

Re: Project Glasswing: Securing critical software for the AI era

#480

Earlier quoted context omitted.

100%, poorly architected software is really difficult to make secure. I think this will extend to AI as well. It will just dial up the complexity of the code until bugs and vulnerabilities start creeping in. At some point, people will have to decide to stop the complexity creep and try to produce minimal software. For any complex project with 100k+ lines of code, the probability that it has some vulnerabilities is ve…

> It doesn't fit into LLM context windows and there aren't enough attention heads to attend to every relevant part. That's for one pass. And that pass can produce a summary of what the code does.

But the summary is likely to summarise out the details which makes the code vulnerable.
Post reply on HN