Live data from Hacker News

Project Glasswing: Securing critical software for the AI era

anthropic.com

861–870 of 921 posts

Re: Project Glasswing: Securing critical software for the AI era

#861

Earlier quoted context omitted.

Is there any actual independent data though, or verification of any of these claims? As it stands this is just a marketing programme for all involved.

What would be the product they're marketing by this campaign?

[dead]

Re: Project Glasswing: Securing critical software for the AI era

#862

Earlier quoted context omitted.

We’ve been suggesting that programmers are going to be replaced by simpler programming languages, gui programming tools, no code tools, low code tools, and now AI. The real big step was when Claude code came out and introduced the agentic loop where it could self validate against tests/linters/tooling, but everything after that had been penned as miraculous when IME it’s a new iteration of the same thing - wild hallu…

> We’ve been suggesting that programmers are going to be replaced by simpler programming languages, gui programming tools, no code tools, low code tools, and now AI. The difference is that it's actually working this time. Non-programmers are writing full apps. Sure, they're simple ones, often just CRUD and UI, but it actually is changing things in a way it never has before. You can't assert something is the same as e…

> The difference is that it's actually working this time. Non-programmers are writing full apps

You can say this about every step along the way. C programmers replaced assembly programmers. Python programmers replaced C programmers. low code tools replaced interal tools teams.

> I'm criticizing you for misrepresenting what claim was made in the first place. No where in your evidence have you shown anyone "walking the claim back".

The claim is that SWES will have their work done by models in 6-12 motnhs. We are _nowhere near_ that 9 months on to it. That's all there is to say it.

> If anything, TFA is claiming evidence of an LLM doing "most" of what SWEs do "end to end" three months ahead of schedule.

TFA based on a model that is so good that it has to be kept from us? from the company that literally can't keep their app up? From the company who shipped an update that didn't launch?

> be my guess, but don't tell falsehoods about the evidence you are presenting.

I mean, I literally posted a quote from the CEO of one of the two major companies saying that SWEs are 6-12 months away from being replaced. This is fantasy talk from a guy who is incentivised to have you believe this. If the claims are that software is changing, and how we're building/deploying software is adapting to that new world then yeah that's fair enough. But the current models, harnesses and tooling are not replacing an SWE unless there's a paradigm shift in the next 3 months. And my point is that we appear t be going backwards, not forwards.

> didn't make GLP1s fake just because they had the same type of hype.

No, GLPs work and that's the difference.

Re: Project Glasswing: Securing critical software for the AI era

#863

> Mythos Preview identified a number of Linux kernel vulnerabilities that allow an adversary to write out-of-bounds (e.g., through a buffer overflow, use-after-free, or double-free vulnerability.) Many of these were remotely-triggerable. However, even after several thousand scans over the repository, because of the Linux kernel’s defense in depth measures Mythos Preview was unable to successfully exploit any of these…

Time to adopt Ada and SPARK.

Re: Project Glasswing: Securing critical software for the AI era

#864

You'd think with this "terrifying" powerful model of theirs they could have a few less red bars on their status page[1], but apparently the hyper-intelligence is only capable of pulling off uber-sophisticated cyber attacks and not making a frontend that doesn't shit itself constantly, curious. [1] https://status.claude.com/

One argument can be made that this is an issue of there simply not being enough compute in the world to meet the demand for claude's LLMs right now, and not really an issue with their infra setup or architecture.

Re: Project Glasswing: Securing critical software for the AI era

#865
post #460

Earlier quoted context omitted.

Everyone’s pretending the suits are going to want to do the prompting. We all know they aren’t.

The suits won't prompt, the model will.

Sounds like the mythical agi I keep hearing about.

Re: Project Glasswing: Securing critical software for the AI era

#866

Earlier quoted context omitted.

Memory tagging has not “crushed hacking” it’s just changed the kinds of exploits that work

That’s underselling it. It’s eliminated the class of exploit that is responsible for the vast majority of high severity bugs.

No they did not

Re: Project Glasswing: Securing critical software for the AI era

#867

OpenAI initially claimed that GPT-2 was too dangerous to release in 2019. How many times will labs repeat the same absurd propaganda?

GPT2 was definitely a risk, just not of the same magnitude. It would have (and did!) make social media bot farms way more convincing and widespread. There was specific worry about that being used to sway elections, which is why they held back the model.

Re: Project Glasswing: Securing critical software for the AI era

#868

Earlier quoted context omitted.

This is obviously just cope (there's a long, strong-form argument for why LLM-agent vulnerability research is plausibly much more potent than fuzzing, but we don't have to reach it because you can dispose of the whole argument by noting that agents can build and drive fuzzers and triage their outputs), but what I'd really like to understand better is why? What's the impetus to come up with these weird rationalization…

You said it yourself. It's cope. That's all it is and all it ever was. https://en.wikipedia.org/wiki/AI_effect Every time an AI does something new, there's a human saying "it's not really doing that something", "it's doing that something in a fake way" or "that something was never important in the first place".

Alright, except that’s not what I was saying. I was just pointing out that LLMs don’t replace fuzzing or static analysis. They complement those techniques. And yes, LLMs may drive those techniques directly, but they often don’t. At least not yet.

Re: Project Glasswing: Securing critical software for the AI era

#869

Now, its very possible that this is Anthropic marketing puffery, but even if it is half true it still represents an incredible advancement in hunting vulnerabilities. It will be interesting to see where this goes. If its actually this good, and Apple and Google apply it to their mobile OS codebases, it could wipe out the commercial spyware industry, forcing them to rely more on hacking humans rather than hacking mobi…

its very possible that this is Anthropic marketing puffery It isn't.

Two possibilities:

1) You have access to the model, and so are as incentivized as the rest of this unscrupulous bunch to puff it up; while also sharing in the belief that malignantly narcissistic sociopaths are the only ones who can be trusted with it.

2) You lack access to the model, and are just doing more PR puffery.

Re: Project Glasswing: Securing critical software for the AI era

#870

I’m sure the new model is a step above the old one but I can’t be the only person who’s getting tired of hearing about how every new iteration is going to spell doom/be a paradigm shift/change the entire tech industry etc. I would honestly go so far as to say the overhype is detrimental to actual measured adoption.

a lot of times people cry wolf for a couple of times before wolf actually comes.

i feel like theres a good chance that this is the actual wolf coming here. cause i was using opus for a lot and it's really good.

Post reply on HN