Live data from Hacker News

Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

alexschapiro.com

101–110 of 301 posts

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#101
post #12

I'm always a bit surprised how long it can take to triage and fix these pretty glaring security vulnerabilities. October 27, 2025 disclosure and November 4, 2025 email confirmation seems like a long time to have their entire client file system exposed. Sure the actual bug ended up being (what I imagine to be) a Is the issue that people aren't checking their security@ email addresses? People are on holiday? These emai…

I'm a bit conflicted about what responsible disclosure should be, but in many cases it seems like these conditions hold:

1) the hack is straightforward to do;

2) it can do a lot of damage (get PII or other confidential info in most cases);

3) downtime of the service wouldn't hurt anyone, especially if we compare it to the risk of the damage.

But, instead of insisting on the immediate shutting down of the affected service, we give companies weeks or months to fix the issue while notifying no one in the process and continuing with business as usual.

I've submitted 3 very easy exploits to 3 different companies the past year and, thankfully, they fixed them in about a week every time. Yet, the exploits were trivial (as I'm not good enough to find the hard ones, I admit). Mostly IDORs, like changing id=123456 to id=1 all the way up to id=123455 and seeing a lot medical data that doesn't belong to me. All 3 cases were medical labs because I had to have some tests done and wanted to see how secure my data was.

Sadly, in all 3 cases I had to send a follow-up e-mail after ~1 week, saying that I'll make the exploit public if they don't fix it ASAP. What happened was, again, in all 3 cases, the exploit was fixed within 1-2 days.

If I'd given them a month, I feel they would've fixed the issue after a month. If I'd given then a year - after a year.

And it's not like there aren't 10 different labs in my city. It's not like online access to results is critical, either. You can get a printed result or call them to write them down. Yes, it would be tedious, but more secure.

So I should've said from the beginning something like:

> I found this trivial exploit that gives me access to medical data of thousands of people. If you don't want it public, shut down your online service until you fix it, because it's highly likely someone else figured it out before me. If you don't, I'll make it public and ruin your reputation.

Now, would I make it public if they don't fix it within a few days? Probably not, but I'm not sure. But shutting down their service until the fix is in seems important. If it was some hard-to-do hack chaining several exploits, including a 0-day, it would be likely that I'd be the first one to find it and it wouldn't be found for a while by someone else afterwards. But ID enumerations? Come on.

So does the standard "responsible disclosure", at least in the scenario I've given (easy to do; not critical if the service is shut down), help the affected parties (the customers) or the businesses? Why should I care about a company worth $X losing $Y if it's their fault?

I think in the future I'll anonymously contact companies with way more strict deadlines if their customers (or others) are in serious risk. I'll lose the ability to brag with my real name, but I can live with it.

As to the other comments talking about how spammed their security@ mail is - that's the cost of doing business. It doesn't seem like a valid excuse to me. Security isn't one of hundreds random things a business should care about. It's one of the most important ones. So just assign more people to review your mail. If you can't, why are you handling people's PII?

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#102

Earlier quoted context omitted.

Do you think the original comment posted by quapster was "slop" equivalent to a copy-paste spam bot? The only spam I see in this chain is the flagged post by electric_muse. It's actually kind of ironic you bring up copy-paste spam bots. Because people fucking love to copy-paste "ai slop" on every comment and article that uses any punctuation rarer than a period.

> Do you think the original comment posted by quapster was "slop" equivalent to a copy-paste spam bot? Yes: the original comment is unequivocally slop that genuinely gives me a headache to read. It's not just "using any punctuation rarer than a period": it's the overuse and mis use of punctuation that serves as a tell. Humans don't needlessly use a colon in every single sentence they write: abusing punctuation like t…

"it's not just x, it's y" is an ai pattern and you just said:

>"It's not just "using any punctuation rarer than a period": it's the overuse and misuse of punctuation that serves as a tell."

So, I'm actually pretty sure you're just copy-pasting my comments into chatgpt to generate troll-slop replies, and I'd rather not converse with obvious ai slop.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#103

Earlier quoted context omitted.

> Do you think the original comment posted by quapster was "slop" equivalent to a copy-paste spam bot? Yes: the original comment is unequivocally slop that genuinely gives me a headache to read. It's not just "using any punctuation rarer than a period": it's the overuse and mis use of punctuation that serves as a tell. Humans don't needlessly use a colon in every single sentence they write: abusing punctuation like t…

"it's not just x, it's y" is an ai pattern and you just said: > "It's not just "using any punctuation rarer than a period": it's the overuse and misuse of punctuation that serves as a tell." So, I'm actually pretty sure you're just copy-pasting my comments into chatgpt to generate troll-slop replies, and I'd rather not converse with obvious ai slop.

Congratulations, you successfully picked up on a pattern when I was intentionally mimicking the tone of the original spambot content to point out how annoying it was. Why are you incapable of doing this with the original spambot comment?

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#104
post #94

It's so great that they allowed him to publish a technical blog post. I once discovered a big vulnerability in a listed consumer tech company -- exposing users' private messages and also allowing to impersonate any user. The company didn't allow me to write a public blogpost.

Why is the control of publication in their hands and not in yours? Shouldn’t you be able to do whatever after disclosing it responsibly?

Presumably they'll threaten to sue you and/or file a criminal complaint, which can be pretty hard to deal with depending on the jurisdiction. At that point you'll probably start asking yourself if it's worth publishing a blog post for some internet points.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#105
post #8

Earlier quoted context omitted.

It's a little hilarious. First, as an organization, do all this cybersecurity theatre, and then create an MCP/LLM wormhole that bypasses it all. All because non-technical folks wave their hands about AI and not understanding the most fundamental reality about LLM software being fundamentally so different than all the software before it that it becomes an unavoidable black hole. I'm also a little pleased I used two sp…

Nitpick, but wormholes and black holes aren't limited to space! (unless you go with the Rick & Morty definition where "there's literally everything in space")

Not a nit pick at all friend, it is even more rabbit holes to explore.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#106

Earlier quoted context omitted.

"it's not just x, it's y" is an ai pattern and you just said: > "It's not just "using any punctuation rarer than a period": it's the overuse and misuse of punctuation that serves as a tell." So, I'm actually pretty sure you're just copy-pasting my comments into chatgpt to generate troll-slop replies, and I'd rather not converse with obvious ai slop.

Congratulations, you successfully picked up on a pattern when I was intentionally mimicking the tone of the original spambot content to point out how annoying it was. Why are you incapable of doing this with the original spambot comment?

I'm not replying to your slop (well, you know, after this one).

Anyways, if you think something is ai, just flag it instead so I don't need to read the word "slop" for the 114th fucking time today.

Thankfully, this time, it was flagged. But I got sucked in to this absolutely meaningless argument because I lack self control.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#107
post #12

I'm always a bit surprised how long it can take to triage and fix these pretty glaring security vulnerabilities. October 27, 2025 disclosure and November 4, 2025 email confirmation seems like a long time to have their entire client file system exposed. Sure the actual bug ended up being (what I imagine to be) a Is the issue that people aren't checking their security@ email addresses? People are on holiday? These emai…

> October 27, 2025 disclosure and November 4, 2025 email confirmation seems like a long time to have their entire client file system exposed

There is always the simple answer, these are lawyers so they are probably scrambling internally to write a response that covers themselves legaly also trying to figure out how fucked they are.

1 week is surprisingly not that slow.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#108

Earlier quoted context omitted.

Congratulations, you successfully picked up on a pattern when I was intentionally mimicking the tone of the original spambot content to point out how annoying it was. Why are you incapable of doing this with the original spambot comment?

I'm not replying to your slop (well, you know, after this one). Anyways, if you think something is ai, just flag it instead so I don't need to read the word "slop" for the 114th fucking time today. Thankfully, this time, it was flagged. But I got sucked in to this absolutely meaningless argument because I lack self control.

Ironically, you were the first person in this thread to use the word "slop". You have become what you hate.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#109

> ... after looking through minified code, which SUCKS to do ... AI tends to be good at un-minifying code.

Legit question: when working on finding security issues, are there any guidelines on what you can send to LLMs/AI?

I got downvoted, so maybe that means someone thinks un-minifying code is not advised for dealing with security issues? But on reflection surely you can just use the 'format code' command in the ide? I am no expert but surely it's ok to use AI to help track down and identify security issues with the usual caveats of 'don't believe it blindly, do your double checking and risk assessing.'

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#110

Earlier quoted context omitted.

I'm not replying to your slop (well, you know, after this one). Anyways, if you think something is ai, just flag it instead so I don't need to read the word "slop" for the 114th fucking time today. Thankfully, this time, it was flagged. But I got sucked in to this absolutely meaningless argument because I lack self control.

Ironically, you were the first person in this thread to use the word "slop". You have become what you hate.

jokes on you, I already hate me, that’s why I spend so much time on HN arguing about nothing

oh shit I’m supposed to be done replying

Post reply on HN