Live data from Hacker News

Anthropic: Expanding Access to Claude for Government

anthropic.com

21–30 of 99 posts

Re: Anthropic: Expanding Access to Claude for Government

#21
post #3

There's no doubt that LLMs massively expand the ability of agencies like the NSA to perform large-scale surveillance at a higher quality. I wonder if Anthropic (or other LLM providers) ever push back or restrict these kinds of use cases? Or is that too risky for them?

NSA should be training their own GPT-4 or better model as we speak and should have been doing it for a long while now. Anything else is borderline incompetence.

Re: Anthropic: Expanding Access to Claude for Government

#22
post #9
post #4

Earlier quoted context omitted.

Will it really though? So far I've seen most of the "revolutionise" claims to be mainly hot air and marketing. It's possible that LLMs will suddenly make a leap in reliability and usability (e.g. much higher context window without corresponding massive increases in memory usage). But I have yet to see it. So far it's great at some specific usecases. Interacting with humans, rewriting or making up text. Summarising. A…

I use LLMs extensively in my field to automate all sorts of tasks. Need to classify a million PDF documents for cheap? Write a prompt and submit a batch job. Need to read 30,000 drilling reports to automatically scan for hazards? Done in 60 minutes. These are tasks that would have taken months of development or millions of dollars in manual effort before. It's not just hype.

Yes but that is one of those niche tasks I meant.

Once again they are selling it like something that's for everyone right now. This is the problem. THe same with the metaverse. It has some really great usecases, but they made it out like next year we would all ditch our phones and work exclusively in a VR headset. Obviously that didn't happen, as the tech was nowhere near that and probably people don't want it either.

Also, if you really need to be sure that those 30.000 drilling reports really didn't contain any hazards, you still have to go through it all yourself. Don't forget LLMs aren't reproducible.

But no, my point was exactly that it's not just hype. There are genuine useful usecases, I totally agree.

As there were for metaverse, and probably even for blockchain (NFT not so sure tho :) I always thought they were really a solution looking for a problem). The key thing about a hype is that they overblow the potential benefits way too much though. I see this happening here once again.

Re: Anthropic: Expanding Access to Claude for Government

#23

I find all of the virtue signalling from AI companies exhausting.

Are you assuming that anyone claiming to be doing good things must be lying? And then getting angry at them both for not doing the good thing and for the lying?

That does sound exhausting.

Re: Anthropic: Expanding Access to Claude for Government

#24
post #9

Earlier quoted context omitted.

I use LLMs extensively in my field to automate all sorts of tasks. Need to classify a million PDF documents for cheap? Write a prompt and submit a batch job. Need to read 30,000 drilling reports to automatically scan for hazards? Done in 60 minutes. These are tasks that would have taken months of development or millions of dollars in manual effort before. It's not just hype.

A genuine question and not meant as a snipe: as hallucinations are an inherent “feature” of LLMs, how can you be sure of the accuracy of the model’s interpretation of those 30,000 drilling report hazards? Or what is the acceptable level of risk?

You have it write a program to analyze it. I think a lot of people fail to understand that you don't always need the LLM to do the thing, have it write a program to do the thing for you.

Re: Anthropic: Expanding Access to Claude for Government

#25
post #9
post #4

Earlier quoted context omitted.

Will it really though? So far I've seen most of the "revolutionise" claims to be mainly hot air and marketing. It's possible that LLMs will suddenly make a leap in reliability and usability (e.g. much higher context window without corresponding massive increases in memory usage). But I have yet to see it. So far it's great at some specific usecases. Interacting with humans, rewriting or making up text. Summarising. A…

I use LLMs extensively in my field to automate all sorts of tasks. Need to classify a million PDF documents for cheap? Write a prompt and submit a batch job. Need to read 30,000 drilling reports to automatically scan for hazards? Done in 60 minutes. These are tasks that would have taken months of development or millions of dollars in manual effort before. It's not just hype.

One thing I've struggled with while applying LLMs to business problems is how others have dealt with identifying and managing system failures.

Let's say some of your drilling reports contain a pattern that indicates balrog activity, which the LLM misses. The legal or insurance context requires you to monitor and address potential balrog activity. How do you plan for these failures?

In almost every case I've seen, the plan is to not have a plan, which is another way of saying that the data doesn't matter so long as no one complains about the results.

Re: Anthropic: Expanding Access to Claude for Government

#26
post #17

Earlier quoted context omitted.

GPT-4 is “holy shit, this actually works, could be better but it’s so good I almost can’t believe it” while GPT-3.5 is “when it works it’s pretty great, just a pity it almost never does”. So I would assume that three letter agencies would love to take something like GPT-4 and fine tune it based on all the data they have about existing terrorists.

I'm still dealing with hallucinations nearly every time I use it.

I get maybe one hallucination per twenty chats with gpt4.

Re: Anthropic: Expanding Access to Claude for Government

#27
post #15
post #9

Earlier quoted context omitted.

I use LLMs extensively in my field to automate all sorts of tasks. Need to classify a million PDF documents for cheap? Write a prompt and submit a batch job. Need to read 30,000 drilling reports to automatically scan for hazards? Done in 60 minutes. These are tasks that would have taken months of development or millions of dollars in manual effort before. It's not just hype.

Boy, I can’t wait for the foundation of my house to disappear because the LLM mis-classified a drilling report as non-hazardous. What’s the deal here with liability and accountability? That’s a serious problem when considering using these for anything other than toy problems.

You don't actually think the LLM is reviewing those 30k documents do you? You tell it to write a program (which is easy to audit) to pull the info from the PDFs or whatever. I don't get why this crowd is so goddamn unimaginative with LLMs.

Re: Anthropic: Expanding Access to Claude for Government

#28
post #9
post #4

Earlier quoted context omitted.

Will it really though? So far I've seen most of the "revolutionise" claims to be mainly hot air and marketing. It's possible that LLMs will suddenly make a leap in reliability and usability (e.g. much higher context window without corresponding massive increases in memory usage). But I have yet to see it. So far it's great at some specific usecases. Interacting with humans, rewriting or making up text. Summarising. A…

I use LLMs extensively in my field to automate all sorts of tasks. Need to classify a million PDF documents for cheap? Write a prompt and submit a batch job. Need to read 30,000 drilling reports to automatically scan for hazards? Done in 60 minutes. These are tasks that would have taken months of development or millions of dollars in manual effort before. It's not just hype.

I think that for low-risk classification tasks and similar, something like an LLM is a great tool, and I can absolutely see it being extremely useful for intelligence work where sifting through stuff is very hard. However, I would not at all trust AI to make actually important decisions independently.

Re: Anthropic: Expanding Access to Claude for Government

#29
post #9
post #4

Earlier quoted context omitted.

Will it really though? So far I've seen most of the "revolutionise" claims to be mainly hot air and marketing. It's possible that LLMs will suddenly make a leap in reliability and usability (e.g. much higher context window without corresponding massive increases in memory usage). But I have yet to see it. So far it's great at some specific usecases. Interacting with humans, rewriting or making up text. Summarising. A…

I use LLMs extensively in my field to automate all sorts of tasks. Need to classify a million PDF documents for cheap? Write a prompt and submit a batch job. Need to read 30,000 drilling reports to automatically scan for hazards? Done in 60 minutes. These are tasks that would have taken months of development or millions of dollars in manual effort before. It's not just hype.

I hope with all the time and money being saved, you're having humans check the results.

Re: Anthropic: Expanding Access to Claude for Government

#30

Earlier quoted context omitted.

A genuine question and not meant as a snipe: as hallucinations are an inherent “feature” of LLMs, how can you be sure of the accuracy of the model’s interpretation of those 30,000 drilling report hazards? Or what is the acceptable level of risk?

You have it write a program to analyze it. I think a lot of people fail to understand that you don't always need the LLM to do the thing, have it write a program to do the thing for you.

Okay, but you still need to debug the program. If your program must give correct results you still need to check the program output against every case. There's no free lunch there.
Post reply on HN