Live data from Hacker News

Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

alexschapiro.com

131–140 of 301 posts

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#131
-The Filevine team was responsive, professional, and took the findings seriously throughout the disclosure process. They acknowledged the severity, worked to remediate the issues, allowed responsible disclosure, and maintained clear communication. This is another great example of how organizations should handle security disclosures.

In the same tenure I think that a professional etical hacker or a curious fellow that is poking around with no harm intent, shouldn't disclose the name of the company that had a security issue if they resolve it professionally.

You can write the same blog post without mentioning that it was Filevine.

If they didn't take care of the incident that's a different story...

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#132

-The Filevine team was responsive, professional, and took the findings seriously throughout the disclosure process. They acknowledged the severity, worked to remediate the issues, allowed responsible disclosure, and maintained clear communication. This is another great example of how organizations should handle security disclosures. In the same tenure I think that a professional etical hacker or a curious fellow that…

Eh, with something this horrendously egregious I think their customers have a right to know how carelessly their data was handled, regardless of the remediation steps taken after disclosure; that aside, who knows how many other AI SaaS vendors might stumble across this article and realize they've made a similarly boneheaded error, and save both themselves and their customers a huge amount of pain . . .

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#133

-The Filevine team was responsive, professional, and took the findings seriously throughout the disclosure process. They acknowledged the severity, worked to remediate the issues, allowed responsible disclosure, and maintained clear communication. This is another great example of how organizations should handle security disclosures. In the same tenure I think that a professional etical hacker or a curious fellow that…

This is a very standard part of responsible disclosure. Hacker finds bugs -> discloses them to the vendor -> (hopefully) the vendor communicates with them and remediates -> both sides publish the technical details. It also helps to demonstrate to the rest of the security world which companies will take reports seriously and which ones won’t, which is very useful information to have.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#134

This might be off topic since we are in topic of AI tool and on HackerNews. I've been pondering a long time how does one build a startup company in domain they are not familiar with but ... Just have this urge to 'crave a pie' in this space. For the longest time, I had this dream of starting or building a 'AI Legal Tech Company' -- big issue is, I don't work in legal space at all. I did some cold reach on lawfirm rel…

I think if you have no domain expertise or unique insight it will be quite hard to find a real pain point to solve, deliver a winning solution, and have the ability to sell it. Not impossible, but very hard. And starting a company is hard enough as it is. So 9/10 times the answer will be to partner with someone who understands the space and pain point, preferably one who has lived it, or find an easier problem to sol…

I would also split the concerns:

1. Compliancy with relevant standards. HIPAA, GDPR, ISO, military, legal, etc. Realistically you're going to outsource this or hire someone who knows how to build it, and then you're going to pay an agency to confirm that you're compliant. You also need to consider whether the incumbent solution is a trust-based solution, like the old "nobody gets fired for buying Intel".

2. Domain expertise is always easier if you have a domain expert. Big companies also outsource market research. They'll go to a firm like GLG, pay for some expert's time or commission a survey.

It seems like table stakes to do some basic research on your own to see what software (or solutions) exist and why everyone uses them, and why competitors failed. That should cost you nothing but time, and maybe expense if you buy some software. In a lot of fields even browsing some forums or Reddit is enough. The difference is if you have a working product that's generic enough to be useful to other domains, but you're not sure. Then you might be able to arrange some sort of quid pro quo like a trial where the partner gets to keep some output/analysis, and you get some real-world testing and feedback.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#135
post #21
post #12

I'm always a bit surprised how long it can take to triage and fix these pretty glaring security vulnerabilities. October 27, 2025 disclosure and November 4, 2025 email confirmation seems like a long time to have their entire client file system exposed. Sure the actual bug ended up being (what I imagine to be) a Is the issue that people aren't checking their security@ email addresses? People are on holiday? These emai…

security@ emails do get a lot of spam. It doesn't get talked about very much unless you're monitoring one yourself, but there's a fairly constant stream of people begging for bug bounty money for things like the Secure flag not being set on a cookie. That said, in my experience this spam is still a few emails a day at the most, I don't think there's any excuse for not immediately patching something like that. I guess…

My favorite one is the "We've identified a security hole in your website"... and I always respond quickly that my website is statically generated, nothing dynamic and immutable on cloudflare pages. For some odd reason, I never hear back from them.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#136
post #3

The first thing that comes to my mind is SOC2 HIPAA and the whole security theater. I am one of the engineers that had to suffer through countless screenshots and forms to get these because they show that you are compliant and safe. While the real impactful things are ignored

SemiAnalysis made this a base requirement for being appropriately ranked on their ClusterMAX report, telling me it is akin to FAA certifications, and then getting hacked themselves for not enforcing simple security controls.

https://jon4hotaisle.substack.com/i/180360455/anatomy-of-the...

It is crazy how this gets perpetuated in the industry as actually having security value, when in reality, it is just a pay-to-play checkbox.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#137

Earlier quoted context omitted.

The users' comment history does read like generic LLM output. Look at the first lines of different comments: > Interesting point about Cranelift! I've been following its development for a while, and it seems like there's always something new popping up. > Interesting point about the color analysis! It kinda reminds me of how album art used to be such a significant part of music culture. > Interesting point about the…

Or maybe these are people who learned from a LLM that English is supposed to sound like this if you want to be permitted to communicate a.k.a. "to be taken into consideration"! Which is wrong and also kinda sucks, but also it sucks and is wrong for a kinda non-obvious reason. Or, bear with me there, maybe things aren't so far downhill yet, these users just learned how English is supposed to sound, from the same place…

Or maybe the 2 month old account posting repetitive comments and using the exact patterns common to AI generated comment is, actually, posting LLM generated content.

> So what if they are? Then they'd just be stupid, futile thoughts leading exactly nowhere.

FYI, spammers love LLM generated posting because it allows them to "season" accounts on sites like Hacker News and Reddit without much effort. Post enough plausible-sounding comments without getting caught and you have another account to use for your upvote army, which is a service you can now sell to desperate marketing people who promised their boss they'd get on the front page of HN. This was already a problem with manual accounts but it took a lot of work to generate the comments and content.

That's the "so what"

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#138
post #95

> November 20, 2025: I followed up to confirm the patch was in place from my end, and informed them of my intention to write a technical blog post. Can that company tell you to cease and desist? How does the law work?

FYI, a "cease and desist" carries the same legal weight as me sending a one-liner saying "Knock it off".

They are strongly worded requests from a legal point of view. The only real message they send is that the sender is serious enough about the issue to have involved a lawyer, unless of course you write it yourself, which is something that literally anyone can do.

If you want to actually force an action, you need a court order of some type.

NB for the actual lawyers: I'm oversimplifying, since they can be used in court to prove that you tried to get the other party to stop, and tried to resolve the issue outside of court.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#139

-The Filevine team was responsive, professional, and took the findings seriously throughout the disclosure process. They acknowledged the severity, worked to remediate the issues, allowed responsible disclosure, and maintained clear communication. This is another great example of how organizations should handle security disclosures. In the same tenure I think that a professional etical hacker or a curious fellow that…

That's not how ethical disclosure works. Both parties should publish and we, the wider tech industry should see this as a good thing both for the hacker and the company that worked with them.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#140
post #48

Who is Margolis, and are they happy that OP publicly announced accessing all their confidential files? Clever work by OP. Surely there is automatic prober tool that already hacked this product?

> Who is Margolis, and are they happy that OP publicly announced accessing all their confidential files?

Google tells me they are a NY law firm specializing in Real Estate and Immigration law. There are other firms with Margolis in the name too. Kinda doesn't matter; see below.

I doubt that they are thrilled to have their name involved in this, but that is covered by the US constitution's protections on free press.

Post reply on HN