Live data from Hacker News

Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

alexschapiro.com

251–260 of 301 posts

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#251

Earlier quoted context omitted.

Prediction: Vibe coding systems will be better at security in 2 years than 90% of devs.

Prediction: it won't. You can't fit every security consideration into the context window.

You may want to read about agentic AI, you can for instance call an LLM multiple times with different security consideration everytime.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#252

Earlier quoted context omitted.

humans used open s3 buckets stuffed with text files of usernames, passwords, addresses, credit card numbers etc long before vibe coding was a thing.

Prediction: Vibe coding systems will be better at security in 2 years than 90% of devs.

Only if they work in a fundamentally different manner. We can't solve that problem the way we are building LLMs now.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#253
post #218

So, 1) a public service, 2) with no authentication, 3) and no encryption? (http only??), 4) sent every single response with a token, 5) giving full admin access to every client's legal documents . This is like a law firm with an open back door, open back window, and all the confidential legal papers sprawled out on the floor. Imagine the potential impact. You're a single mother, fighting for custody of your kids. You…

This is HN. We understood exactly what “exposed … confidential files” meant before reading your overly dramatic scenario. As overdone as it is, it’s not even realistic. A likely single mother is likely tiny potatoes in comparison to deep-pocketed legal firms or large corporations. The story is an example of the market self-correcting , but out comes this “building code” hobby horse anyway. All a software “building co…

[dead]

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#254

Earlier quoted context omitted.

90% of human devs can fit more than 3-5 files into their short-term memory. They also know not to, say, temporarily disable auth to be able to look at the changes they've made on a page hidden behind auth, which is what I observed Gemini 3 Pro doing just yesterday.

Ok, and that’s your prediction for 2 years from now? It’d be quite remarkable if humans had a bigger short term memory than LLMs in 2 years. Or that the kind of dumb security mistakes LLMs make today don’t trigger major, rapid improvements.

Do you understand what the term "context window" means? Have you ever tried using an LLM to program anything even remotely complex? Have you observed how the quality of the output drastically reduces the longer the coversation gets?

That's what makes it bad at security. It cannot comprehend more than a floppy drive worth of data before it reverts to absolute gibberish.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#256
post #147

I work for a finance firm and everyone is wondering why we can store reams of client data with SaaS Company X, but not upload a trust document or tax return to AI SaaS Company Y. My argument is we're in the Wild West with AI and this stuff is being built so fast with so many evolving tools that corners are being cut even when they don't realize it. This article demonstrates that, but it does sort of beg the question…

While the FileVine service is indeed a Legal AI tool, I don't see the connection between this particular blunder and AI itself. It sure seems like any company with an inexperienced development team and thoughtless security posture could build a system with the same issues. Specifically, it does not appear that AI is invoked in any way at the search endpoint - it is clearly piping results from some Box API.

There is none. Filevine is not even an "AI" company. They are a pretty standard SaaS that has some AI features nowadays. But the hive mind needs its food, and AI bad as we all know.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#257
post #147

I work for a finance firm and everyone is wondering why we can store reams of client data with SaaS Company X, but not upload a trust document or tax return to AI SaaS Company Y. My argument is we're in the Wild West with AI and this stuff is being built so fast with so many evolving tools that corners are being cut even when they don't realize it. This article demonstrates that, but it does sort of beg the question…

While the FileVine service is indeed a Legal AI tool, I don't see the connection between this particular blunder and AI itself. It sure seems like any company with an inexperienced development team and thoughtless security posture could build a system with the same issues. Specifically, it does not appear that AI is invoked in any way at the search endpoint - it is clearly piping results from some Box API.

> any company with an inexperienced development team and thoughtless security posture

Point out one (1) "AI product" company that isn't described accurately by that sentence

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#258
post #206

Earlier quoted context omitted.

> it's almost as if there's more stuff we do than just write code.. Yes, but adding these common sense considerations is actually something LLMs can already do reasonably well.

If you explicitly request it which means you need to know about it.

OpenAI can put that in the system prompt for their CTO-as-a-service once, and then forget about it.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#259

Earlier quoted context omitted.

but google told me everyone can vibe code apps now and software engineers should count their days... it's almost as if there's more stuff we do than just write code...

humans used open s3 buckets stuffed with text files of usernames, passwords, addresses, credit card numbers etc long before vibe coding was a thing.

And those humans would be looking for a new job or face other consequences. An AI model can merrily do this with zero consequences because no meaningful consequences can be visited upon it.

Just like if any human employee publicly sexually harassed his female CEO, he'd be out of a job and would find it very hard to find a new one. But Grok can do it and it's the CEO who ends up quitting.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#260
post #223

If they have a billion dollar valuation, this fairly basic (and irresponsible) vulnerability could have cost them a billion dollars. If someone with malice had been in your shoes, in that industry, this probably wouldn't have been recoverable. Imagine a firm's entire client communications and discovery posted online. They should have given you some money.

Who says they didn't give him money?

[deleted]
Post reply on HN