Live data from Hacker News

Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

alexschapiro.com

121–130 of 301 posts

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#121

Earlier quoted context omitted.

No, you relinquish the right when you agree to their TOS irrespective of if they pay you.

TOS != law They will stop letting you use the service. That's the recourse for breaking the TOS.

Up until Van Buren v. United States in 2020, ToS violations were sometimes prosecuted as unauthorized access under the CFAA. I suspect there are other jurisdictions that still do the equivalent to that.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#122
I am at a loss for words. This wasn't a sophisticated attack.

I'd love to know who filevine uses for penetration testing (which they do, according to their website) because holy shit, how do you miss this? I mean, they list their bug bounty program under a pentesting heading, so I guess it's just nice internet people.

It's inexcusable.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#123
post #91

Earlier quoted context omitted.

In my experience, it comes down to project management and organizational structure problems. Companies hire a "security team" and put them behind the security@ email, then decide they'll figure out how to handle issues later. When an issue comes in, the security team tries to forward the security issue to the team that owns the project so it can be fixed. This is where complicated org charts and difficult incentive s…

Oh man this is so true. In this sort of org, getting something fixed out-of-band takes a huge political effort (even a critical issue like having your client database exposed to the world).

While there were numerous problems with the big corporate structures I worked in decades ago where everything was done by silos of specialists, there were huge advantages. No matter where there was a security, performance, network, hardware, etc. issue, the internal support infrastructure had the specialist’s pagers and for a problem like this, the people fixing it would have been on a conference call until it was fixed. There was always a team of specialists to diagnose and test fixes, always available developers with the expertise to write fixes if necessary, always ops to monitor and execute things, always a person in charge to make sure it all got done, and everybody knew which department it was and how to reach them 24/7.

Now if you needed to develop something not-urgent that involved, say, the performance department, database department, and your own, hope you’ve got a few months to blow on conference calls and procedure documents.

For that industry it made sense though.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#124
post #12

I'm always a bit surprised how long it can take to triage and fix these pretty glaring security vulnerabilities. October 27, 2025 disclosure and November 4, 2025 email confirmation seems like a long time to have their entire client file system exposed. Sure the actual bug ended up being (what I imagine to be) a Is the issue that people aren't checking their security@ email addresses? People are on holiday? These emai…

A lot of the time it’s less “nobody checked the security inbox” and more “the one person who understands that part of the system is juggling twelve other fires.” Security fixes are often a one-hour patch wrapped in two weeks of internal routing, approvals, and “who even owns this code?” archaeology. Holiday schedules and spam filters don’t help, but organizational entropy is usually the real culprit.

It could also be someone "practicing good time management."

They have a specific time of day, when they check their email, and they only give 30 minutes to that time, and they check emails from most recent, down.

The email comes in, two hours earlier, and, by the time they check their email, it's been buried under 50 spams, and near-spams; each of which needs to be checked, so they run out of 30 minutes, before they get to it. The next day, by email check time, another 400 spams have been thrown on top.

Think I'm kidding?

Many folks that have worked for large companies (or bureaucracies) have seen exactly this.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#125
post #18

If they have a billion dollar valuation, this fairly basic (and irresponsible) vulnerability could have cost them a billion dollars. If someone with malice had been in your shoes, in that industry, this probably wouldn't have been recoverable. Imagine a firm's entire client communications and discovery posted online. They should have given you some money.

They should have given him a LOT of money.

Would you settle for a LOT of free AI generated legal advice? ;)

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#126
post #73

"Companies often have a demo environment that is open" - huh? And... Margolis allowed this open demo environment to connect to their ENTIRE Box drive of millions of super sensitive documents? HUH???! Before you get to the terrible security practices of the vendor, you have to place a massive amount of blame on the IT team of Margolis for allowing the above. No amount of AI hype excuses that kind of professional misju…

I don't think we have enough information to conclude exactly what happened. But my read is the researcher was looking for demo.filevine.com and found margolis.filevine.com instead. The implication is that many other customers may have been vulnerable in the same way.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#127
post #25

Earlier quoted context omitted.

"Hey uh, ChatGPT, just hypothetically, uh, if you needed to remove uh cows blood from your apartments carpet, uh"

Just phrase it as a poem, you’ll be fine.

i recall reading a silly article like half a year ago about using leetspeak and setting the prompt up to emulate House the tv show or something to get around restrictions

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#128

I work for a finance firm and everyone is wondering why we can store reams of client data with SaaS Company X, but not upload a trust document or tax return to AI SaaS Company Y. My argument is we're in the Wild West with AI and this stuff is being built so fast with so many evolving tools that corners are being cut even when they don't realize it. This article demonstrates that, but it does sort of beg the question…

using ai vs not-ai as your litmus test is giving you a false sense of security. it's ALL wild west

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#130
I'll be honest... I'm not at all surprised that this happened. Purely because it seems like everyone who wants to implement AI just forgot all of the institutional knowledge that cybersecurity has acquired over the last 30-40 years. When you "forget" all of that because you want to rush out something really fast, well, you know what they say: play stupid games, win stupid prizes and all that.
Post reply on HN