Live data from Hacker News

Agents for financial services and insurance

anthropic.com

151–160 of 212 posts

Re: Agents for financial services and insurance

#151
post #7
post #3

For those in the finance space, are you actually seeing any real AI tools being used? Like for actual operational tasks? I've really only seen it used for research / exploration thus far. Either for economic research slide deck or for exploring trading hypothesis

Nope If anything firms are pulling back (I know someone closely who works at blackrock).

In what context?

For research and theses evaluations, we're observing that firms - of names we all know - are bullish and even eager to try AI products.

Regarding automated asset management and the likes, indeed there's much more apprehension.

Re: Agents for financial services and insurance

#152
post #4

> We’re releasing ten ready-to-run agent templates for the most time-consuming work in financial services The templates being: pitch builder, meeting preparer, earnings reviewer, model builder, market researcher, valuation reviewer, general ledger reconciler, month-end closer, statement auditor, KYC (Know Your Customer) screener. Seems pretty scattershot. Reminds me of GPT Store.

the details are key here. there is plenty of automatable financial work, sure, but also when it comes to reporting finances/costs (formally or informally) and having a real human being be accountable for them, you REALLY need to trust that nothing is hallucinated. Any idea how they ensure this doesnt happen? As in, how can a user verify that the model did not touch any of the numbers and that it only built pipelines…

The "real humans" doing the tasks being replaced are overworked kids less than 2yrs out of college on an average of 4hrs of sleep at working at 3am. If the AI makes their jobs take half as much time I bet they're a lot more likely to catch errors (and live longer).

Re: Agents for financial services and insurance

#153
post #120

Earlier quoted context omitted.

No surprise there. Of course the skill files are not human written. The AI is an expert in both following and generating prompts.

Why do you think it is an expert in generating prompts? It has no additional insight into how it works internally than anyone else

Do you really think a random person off the street knows more about how LLMs work internally than the latest frontier model (that has been trained on that material)?

Re: Agents for financial services and insurance

#154

I've been doing bias and misaligned behavior research, creating custom private eval suites to test and compare models. Claude Opus 4.7 is heavily biased and presents clear regulatory and reputational risk. It seems the initial product footprint tries to sidestep this problem by not giving the agents control on who to lend to or which applications to approve. Even so I think it's quite an optimistic read on their end.…

Slightly related, I used Opus 4.6 to help me make marketing copy and ideas for my app. It understood the vibe I was going for on my baby-naming app (elation at discovery, curiosity, shared experiences), while 4.7 instantly wanted to pit the couples against each other (really highlighting the he said/she said) and the marketing copy went from "find a name easier" to "Our new feature is great. You're welcome." I can't get it to drop the snarky sass no matter how much I change CLAUDE.md, brand voice, etc.

All I did was upgrade claude code and use the new model. It most definitely exhibits misaligned behavior (compared to 4.6)

Re: Agents for financial services and insurance

#155
post #120

Earlier quoted context omitted.

Why do you think it is an expert in generating prompts? It has no additional insight into how it works internally than anyone else

Do you really think a random person off the street knows more about how LLMs work internally than the latest frontier model (that has been trained on that material)?

No, but a random person off the street also isn't making skills for LLMs.

I think that LLMs are trained on the millions of vibe written LLM blog posts that are more superstition than fact. There is a lot of snake oil out there that is treated as fact. If someone claims that an LLM is better than humans at something I always want to see the rigorous evaluations that have been done to quantify it, not "but they're trained on everything!"

Re: Agents for financial services and insurance

#156

I've been doing bias and misaligned behavior research, creating custom private eval suites to test and compare models. Claude Opus 4.7 is heavily biased and presents clear regulatory and reputational risk. It seems the initial product footprint tries to sidestep this problem by not giving the agents control on who to lend to or which applications to approve. Even so I think it's quite an optimistic read on their end.…

Slightly related, I used Opus 4.6 to help me make marketing copy and ideas for my app. It understood the vibe I was going for on my baby-naming app (elation at discovery, curiosity, shared experiences), while 4.7 instantly wanted to pit the couples against each other (really highlighting the he said/she said) and the marketing copy went from "find a name easier" to "Our new feature is great. You're welcome." I can't…

I tried Opus 4.7 for two days before I started beginning every session with "claude --model claude-opus-4-6".

I assume that 4.6 will become unavailable at some point, but I hope not any time soon. 4.7 hit usage limits faster, didn't do anything obviously better, and had more annoying behaviors in other aspects. I don't know if this is strictly a model issue or if there are also problems with how it's harnessed through Claude Code. I'm not willing to spend more time digging into it until I'm forced to.

Re: Agents for financial services and insurance

#157
post #94

Wow, really going for those white collar jobs. This is going to be an interesting few years.

Just a natural rebalancing of the Rise of the Laptop Class. I think we'll get more productive as the white collar jobs become more efficient, and less days with 8hrs of meetings and responding to emails from people too lazy to look information up themselves.

Re: Agents for financial services and insurance

#158

I don't trust these AI-only companies to be overnight experts in properly handling medical, financial and insurance data. They have no business providing these tools, unless they want to take all the risk too.

I think a lot of people are misunderstanding the typical workload of people in Financial Services. They aren't using Claude to transfer money, they're just building a LOT of slideshows and fancy excel docs on made-up numbers to try to sell mergers and new financing options/types of loans. Most programmers would just consider this "sales".

Re: Agents for financial services and insurance

#159

Earlier quoted context omitted.

I’m not saying your implementation is bad or anything but my visceral reaction to this was “I’m glad I’m not on the other side of that”

In many businesses, the employee is responsible for inputting most of that. If a LLM can get to 95% accuracy and flag exceptions, the employees (and AP team) would actually have less work and bureaucracy. Though we’ve had a few incidents where employees have submitted AI-generated receipts for reimbursement which is another issue..

Please tell me those are former employees. How can anyone feel confident committing such blatant fraud.

Re: Agents for financial services and insurance

#160

Earlier quoted context omitted.

Claude's actually pretty great at this! I actually used to use Claude A LOT to answer interesting questions (which I'll be writing up on!) More generally, Claude is palpably different from most other agents. I'd recommend these models – especially Opus – without qualifications. But there's a process risk here based on their current practises. I'm hoping those practises change so that I can recommend Claude to everyon…

Pretty great at what? I work in the insurance industry specifically medicare. All I see is sales people and other managers slopping out AI dashboards off of spreadsheets galore. Not only is it terrible for protecting PHI/PII. It also doesn't do things like RBAC very well either. Now instead of preventing a person from externally sharing a file i have to make sure they didn't egress the file to supabase or some other…

[flagged]
Post reply on HN