Live data from Hacker News

Discovering Cryptographic Weaknesses with Claude

anthropic.com

61–70 of 199 posts

Re: Discovering Cryptographic Weaknesses with Claude

#61

When high quality effort is applied to a tool, such as AES or the linux kernel, we intuit that it "hardens" the tool. That is, it makes the tool more correct, more resilient, less assailable, etc. Similarly, when effort is applied to an open problem, such as the Riemann hypothesis or P v NP, without progress, it "hardens" the problem: it makes the problem feel more daunting to whoever takes a stab at it next. Andrew…

I wouldn't worry about too many mathematicians adopting the "even AI couldn't solve it" attitude.

Business folks riding the hype train? Maybe.

Re: Discovering Cryptographic Weaknesses with Claude

#62
I'm looking forward to seeing similar work on SHA-256. It would be fascinating if AI could discover previously unknown weaknesses in reduced-round variants.

It would also be interesting whether AI could discover new algorithmic optimizations for SHA-256 similar in spirit to AsicBoost[1].

[1] https://arxiv.org/pdf/1604.00575

Re: Discovering Cryptographic Weaknesses with Claude

#63

When high quality effort is applied to a tool, such as AES or the linux kernel, we intuit that it "hardens" the tool. That is, it makes the tool more correct, more resilient, less assailable, etc. Similarly, when effort is applied to an open problem, such as the Riemann hypothesis or P v NP, without progress, it "hardens" the problem: it makes the problem feel more daunting to whoever takes a stab at it next. Andrew…

This is a problem that will solve itself, people will continue to work on the problems that AI fails at, likely by telling AI the approaches they want AI to take.

Re: Discovering Cryptographic Weaknesses with Claude

#64

Earlier quoted context omitted.

Have you tried using Wispr or Willow (or any one of a thousand alternatives?) A little odd at first but absolutely amazing for the purpose of piling context into an LLM.

It's so bizarre to me that people want to do this. Can't you type faster than you speak? Doesn't your speaking inhibit your thinking? Aren't you self-conscious talking out loud? How are our experiences so different?

> Can't you type faster than you speak? Doesn't your speaking inhibit your thinking?

For me personally: no, I speak faster than I type; and speaking actually helps me get more ideas compared to typing.

(Not sure if that’s due to having no typing speed barrier, or maybe because speaking activates different parts of the brain.)

Once you get over the feeling of self-consciousness, it’s a great way. I even go on short walks sometimes and mumble to my phone to prepare some long prompt. Thinking works even better, when walking outside :-)

Re: Discovering Cryptographic Weaknesses with Claude

#65
post #22

I find that some of my friends and acquaintances have gotten obsessed with prompting style, "prompt engineering", which skills to use, which skills to build, "context engineering", and a billion other variations on "how to write smart things so the model does good". Friends, look at the prompts that Anthropic's own people are putting into the machine: > A few hours after the first message, we found that Claude was st…

To quote - what I found to be - an absolute zinger from another trending thread on here just 6 hours ago:

> Typical users run software written by atypical users.

https://news.ycombinator.com/item?id=49084936

This extends to everything. Anthropic has a few thousand engineers, but millions of (also engineer) users. Entire business can be built on niches that are at most a few week pet project for a team there, that can inevitably and significantly outperform them, despite being the people behind the thing.

I'm sure I'm not the only one here who jumped into this whole agentic stuff, built some tooling to make things comfy, only to see that tooling all be increasingly introduced as prim and proper features in the various harnesses weeks later.

Re: Discovering Cryptographic Weaknesses with Claude

#66

Earlier quoted context omitted.

It's so bizarre to me that people want to do this. Can't you type faster than you speak? Doesn't your speaking inhibit your thinking? Aren't you self-conscious talking out loud? How are our experiences so different?

These are all interrelated points and the sibling comment is correct: it's a skill. The key thing with these voice systems is that you do not need to edit anything. You can literally just stream of consciousness into them, no editing, include the backtracking, the live-revisions, etc., and it will actually all produce vastly better context for the LLM than the written thing you took even 30 seconds to edit for clarit…

You don't need to edit anything when you are typing either. You don't even need to worry about spelling or typos.

Re: Discovering Cryptographic Weaknesses with Claude

#67
post #22

I find that some of my friends and acquaintances have gotten obsessed with prompting style, "prompt engineering", which skills to use, which skills to build, "context engineering", and a billion other variations on "how to write smart things so the model does good". Friends, look at the prompts that Anthropic's own people are putting into the machine: > A few hours after the first message, we found that Claude was st…

Yeah, in fact you can actually make the model perform worst. You should allow the model to "think for itself" instead of pushing your reasoning into the prompt. You should give it simple prompt and steer it along the way. Skills, CLAUDE.md/AGENTS.md should only ever be used if the model struggle at something or doesn't know how to use something. Vast majority of project should never need a skill or CLAUDE.md. If you…

You need agents.md and similar for indications about stuff that is not in the code itself. There are plenty of use cases for that, the alternative is have the LLM guess the most probable solution, which may be correct but may also be wrong.

Re: Discovering Cryptographic Weaknesses with Claude

#68
post #22

I find that some of my friends and acquaintances have gotten obsessed with prompting style, "prompt engineering", which skills to use, which skills to build, "context engineering", and a billion other variations on "how to write smart things so the model does good". Friends, look at the prompts that Anthropic's own people are putting into the machine: > A few hours after the first message, we found that Claude was st…

It's counter to what sci-fi taught us using AI would be like. We never thought we'd have to feed it words of encouragement, we expected it to act more mechanically, like the computer interfaces we have been using, but here we are. It's kind of quaint, and kind of endearing.

C-3PO was neurotic and needed lots of reassurance.

Re: Discovering Cryptographic Weaknesses with Claude

#69
post #5

I can already picture the faces of national security directors everywhere. "The attacks described in these two papers are the strongest attacks we have found to date. We are sharing them after a period of consultation with US government and industry leaders. But as we develop increasingly powerful cryptanalytic results, it would be prudent to consider how researchers should react if a language model were to discover…

> And a veiled pitch to real cryptanalysis researchers: "Researchers at Anthropic then spent several hundred hours learning enough cryptography research to validate the model’s claim"

Many of those researchers, particularly the primary researchers and the individual(s) driving the prompts behind these big stories, have very advanced math degrees and experience. What this shows more than anything is how ML can augment expertise, the searching of solution spaces, and the connecting of dots between existing almost-there research.

But also what's left out is all the time wasted pursuing dead-ends. There's an obvious publication bias at play here, though we can't know how extreme without transparency.

Re: Discovering Cryptographic Weaknesses with Claude

#70
post #18

“Each of the results cost roughly $100,000 in API cost to develop.” And “Over the course of a week, one Anthropic researcher worked together with Claude to develop the HAWK attack, and another researcher built a scaffold4 that allowed Claude to fully autonomously discover the AES attack.” Spending $100k in tokens in a week is an impressive feat even with massive parallelization. I suspect the TPS their internal folks…

That's not really that ridiculous. Looking at my ChatGPT stats my biggest day of token usage was 1B tokens (seeing how far Sol Ultra could go on a difficult problem with a quantitative goal and eval harness that it could run on it's own that allowed it to keep going until it succeeded). I blew through my $100 subscription usage in that one day, but with a 80/20 token blend that's $10k in API billing. So, $70k in a we…

How much was valuable the output you got during the day?

Is it at least comparable to the $10k of cost?

Post reply on HN