Live data from Hacker News

Expanding Project Glasswing

anthropic.com

21–30 of 261 posts

Re: Expanding Project Glasswing

#22
post #12

Earlier quoted context omitted.

Social engineering as a problem goes away when anybody can get a model to do it for them for $5. It stops being possible, it's really the bank's problem when they can't have a minimum wage call center or a robot responsible for people's data.

Yes. There will be a few high-profile incidents, and then institutions will be forced to stop performing administrative actions based on people’s word.

This outcome is massively detrimental to humanity at large. By eliminating the human factor from support, you make it impossible to get support in edge cases that fall outside of the pre-planned bureacratic process. Everyone already hates that Google can arbitrarily ban anybody they please with no way to get in contact with a human, and you want to extend that to banks in control of people's life savings?

Re: Expanding Project Glasswing

#23
post #16
post #6

GPT-5.5-Cyber has already at least hit if not surpassed Mythos capability in cyber tasks. The only reason they're holding back is because once its out everyone would realize that its capabilities were a step change in March, but are not anymore, yet it costs significantly more and is much slower.

So you believe one marketing department more than the other?

The brits have a step-based benchmark that they use for this - https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5...

They seem pretty close, in both average and "best run" scores. And, in a highly verifiable domain, "best run" or pass@n is what you're looking for.

Re: Expanding Project Glasswing

#25
post #13

> The organizations in this new group are based in more than 15 countries I mean most nasdaq tech companies would be in 13+ countries, why are they writing this like it's a big number, is hilariously small?

I assume they're using a more candid definition where they're not counting all the countries a company may be based, but rather the primary country they're based in.

I don't think they're trying to flex this as a large number. They don't want to give an exact number, as that may change etc / is fuzzy, but also want to give you an idea of the scale.

They say "In the future, we intend to expand our geographical reach much further". I imagine this commentary is somewhat related to the concerns that AI will create an even worse "global underclass". AI developments are first accessible to Americans, then allies, and then later the whole world.

Re: Expanding Project Glasswing

#26
post #20

Is there any evidence Mythos is qualitatively better than the Opus 4.x? I'm afraid that the usual mantra that "we just need more scale" that worked well for attracting investments, is not working anymore - bigger models provide marginal improvements while naturally get much more expensive to run. Is this why both Anthropic and OpenAI are rushing for IPOs this year?

From what I've read so far it's less about Mythos being much better at tasks in isolation.

Security wise, it's about being able to find and chain multiple vulnerabilities to actually create viable exploits.

So I would imagine that if you were using it for regular software development you may not feel that it's that different unless used in a particular way?

Re: Expanding Project Glasswing

#27
Step 1: claim you created a tool so dangerous you can't release it

Step2: offer to test it, but only for the biggest companies in the world

Step 3: onboard those big players on your tooling and product

Step 4: profit

This is genius.

Re: Expanding Project Glasswing

#28
They keep writing like they stand to profit from this or something. Too many “coulds” in there for me too, this could be an amazing advancement and it could be nothing… normally we look at data and last headline I saw was 25 “high” vulnerabilities at the cost of $1 million in tokens.

No comparison to human teams, and I’m sure that $1 million in tokens was used by humans, in a team. So like most AI, they’ve developed a tool that capable people can use to be better, but unlike most tools, they’re claiming this to be outright magic. The magic is the hype train.

Re: Expanding Project Glasswing

#29
post #12

Earlier quoted context omitted.

Yes. There will be a few high-profile incidents, and then institutions will be forced to stop performing administrative actions based on people’s word.

This outcome is massively detrimental to humanity at large. By eliminating the human factor from support, you make it impossible to get support in edge cases that fall outside of the pre-planned bureacratic process. Everyone already hates that Google can arbitrarily ban anybody they please with no way to get in contact with a human, and you want to extend that to banks in control of people's life savings?

> Everyone already hates that Google can arbitrarily ban people

Yet they’re still the predominate search engine, sadly the concerns of the few don’t interest monopolistic profit seekers without forced regulations, think how airlines are legally required to give refunds for delayed flights, there’s a reason it required legislation

Re: Expanding Project Glasswing

#30
post #12

Earlier quoted context omitted.

Yes. There will be a few high-profile incidents, and then institutions will be forced to stop performing administrative actions based on people’s word.

This outcome is massively detrimental to humanity at large. By eliminating the human factor from support, you make it impossible to get support in edge cases that fall outside of the pre-planned bureacratic process. Everyone already hates that Google can arbitrarily ban anybody they please with no way to get in contact with a human, and you want to extend that to banks in control of people's life savings?

I don't think anyone is saying that. You will just need to be authenticated before giving any commands to the bank. Maybe some type of TOTP that you can use over the phone or in person.
Post reply on HN