Live data from Hacker News

Anthropic: Expanding Access to Claude for Government

anthropic.com

51–60 of 99 posts

Re: Anthropic: Expanding Access to Claude for Government

#51

Going forward be very very wary of inputting sensitive information in Anthropic, OpenAI products, especially if you work for a foreign government, corporation. Listen to Edward Snowden. This guy is not fucking around.

Edward Snowden presented sales decks for half baked programs as if they were fully realized, for shock value. To sell his narrative. His claims, like everyone elses, should be approached critically.

Re: Anthropic: Expanding Access to Claude for Government

#54

Earlier quoted context omitted.

Okay, but you still need to debug the program. If your program must give correct results you still need to check the program output against every case. There's no free lunch there.

Speaking generally: The program doesn't always have to give correct results. The program just needs to reduce 30k documents down to 200 documents for human review. You're comparing LLMs to a hypothetical alternative where a human reviews all 30k documents in detail. But the real alternative is often just a worse quality sieve where more errors blunder their way through the existing flawed processes. LLMs can improve…

The epistemology problem never goes away. How should I have any confidence that it's correctly flagging things for review? I need to go through 28800 documents to see if it missed anything.

You're right, I am comparing it to that alternative. There are fields and applications where this is necessary. I do not know if drilling reports are one of them. If you can tolerate a large false negative rate then great. But if you need to be catching 99.99% of problems then IMO you should at least be able to show your work. Taking black box output and throwing it over the wall sounds so sketchy in engineering contexts.

Re: Anthropic: Expanding Access to Claude for Government

#55
post #42
post #21

Earlier quoted context omitted.

NSA should be training their own GPT-4 or better model as we speak and should have been doing it for a long while now. Anything else is borderline incompetence.

And given the volume of data they likely sift through, I'd also expect them to want very small, high-throughput models for identifying targets for larger models to examine. On the flip side, LLMs must give the NSA a new challenge: a flood of garbage text generated by no-one in particular. Perhaps there will be more effort to put surveillance directly on-device as tapping networks yields more noise.

On the grasping side, they are probably in the best position to train a GPT-5, given the amount and type of data they're presumed to have.

Re: Anthropic: Expanding Access to Claude for Government

#56
post #48
post #26

Earlier quoted context omitted.

I get maybe one hallucination per twenty chats with gpt4.

I haven't tried more than a handful of queries, but I think I've gotten 100% rate of hallucination or generic useless response to specific question.

Can I try your question? Just curious.

Re: Anthropic: Expanding Access to Claude for Government

#57

Earlier quoted context omitted.

One thing I've struggled with while applying LLMs to business problems is how others have dealt with identifying and managing system failures. Let's say some of your drilling reports contain a pattern that indicates balrog activity, which the LLM misses. The legal or insurance context requires you to monitor and address potential balrog activity. How do you plan for these failures? In almost every case I've seen, the…

Same way you manage human failures?

The way we manage human failures are with rules, checklists, and accountability. LLMs struggle with all of these, and I get the sense that spending 6mos to develop long lists of rules isn't what the parent comment has in mind with "just write a prompt"

Re: Anthropic: Expanding Access to Claude for Government

#58

Earlier quoted context omitted.

A genuine question and not meant as a snipe: as hallucinations are an inherent “feature” of LLMs, how can you be sure of the accuracy of the model’s interpretation of those 30,000 drilling report hazards? Or what is the acceptable level of risk?

How can you be sure with humans doing the work?

That's where the law comes in. You can prosecute a human for negligence. What about an AI?

Re: Anthropic: Expanding Access to Claude for Government

#59
post #21
post #3

There's no doubt that LLMs massively expand the ability of agencies like the NSA to perform large-scale surveillance at a higher quality. I wonder if Anthropic (or other LLM providers) ever push back or restrict these kinds of use cases? Or is that too risky for them?

NSA should be training their own GPT-4 or better model as we speak and should have been doing it for a long while now. Anything else is borderline incompetence.

NSA can't hire the right talent capable of producing that product for the same reason they have trouble finding white-hat security people to hire: You can't work for the government and do drugs in your personal time. Enough of the pie of elite researchers are in to wacky mind-bending that it's a real recruitment problem.

Re: Anthropic: Expanding Access to Claude for Government

#60
post #42
post #21

Earlier quoted context omitted.

NSA should be training their own GPT-4 or better model as we speak and should have been doing it for a long while now. Anything else is borderline incompetence.

And given the volume of data they likely sift through, I'd also expect them to want very small, high-throughput models for identifying targets for larger models to examine. On the flip side, LLMs must give the NSA a new challenge: a flood of garbage text generated by no-one in particular. Perhaps there will be more effort to put surveillance directly on-device as tapping networks yields more noise.

I’d expect they’re using huge models to train many small ones, one for each threat actor. Those small models could decide whether their actor is detected, or it’s time to slot in a different one.
Post reply on HN