Live data from Hacker News

Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

alexschapiro.com

81–90 of 301 posts

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#81
post #8
post #4

[flagged]

It's a little hilarious. First, as an organization, do all this cybersecurity theatre, and then create an MCP/LLM wormhole that bypasses it all. All because non-technical folks wave their hands about AI and not understanding the most fundamental reality about LLM software being fundamentally so different than all the software before it that it becomes an unavoidable black hole. I'm also a little pleased I used two sp…

Maybe this is the key takeaway of GenAI: that some access to data, even partially hallucinated data, is better than the hoops that the security theatre puts in place that prevents average Joe doing their job.

This might just be a golden age for getting access to the data you need for getting the job done.

Next security will catch up and there'll be a good balance between access and control.

Then, as always security goes to far and nobody can get anything done.

It's a tale as old as computer security.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#82
post #11

Earlier quoted context omitted.

That comment didn't read like AI generated content to me. It made useful points and explained them well. I would not expect even the best of the current batch of LLMs to produce an argument that coherent. This sentence in particular seems outside of what an LLM that was fed the linked article might produce: > What's wild is that nothing here is exotic: subdomain enumeration, unauthenticated API, over-privileged token…

The users' comment history does read like generic LLM output. Look at the first lines of different comments: > Interesting point about Cranelift! I've been following its development for a while, and it seems like there's always something new popping up. > Interesting point about the color analysis! It kinda reminds me of how album art used to be such a significant part of music culture. > Interesting point about the…

Or maybe these are people who learned from a LLM that English is supposed to sound like this if you want to be permitted to communicate a.k.a. "to be taken into consideration"! Which is wrong and also kinda sucks, but also it sucks and is wrong for a kinda non-obvious reason.

Or, bear with me there, maybe things aren't so far downhill yet, these users just learned how English is supposed to sound, from the same place where the LLMs learned how English is supposed to sound! Which is just the Internet.

AI hype is already ridiculous; the whole "are you using an AI to write your posts for you" paranoia is even more absurd. So what if they are? Then they'd just be stupid, futile thoughts leading exactly nowhere. Just like most non-AI-generated thoughts, except perhaps the one which leads to the fridge.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#83
post #27
post #21

Earlier quoted context omitted.

security@ emails do get a lot of spam. It doesn't get talked about very much unless you're monitoring one yourself, but there's a fairly constant stream of people begging for bug bounty money for things like the Secure flag not being set on a cookie. That said, in my experience this spam is still a few emails a day at the most, I don't think there's any excuse for not immediately patching something like that. I guess…

This. There is so much spam from random people about meaningless issues in our docs. AI has made the problem worse. Determining the meaningful from the meaningless is a full time job.

This is where “managed” bug bounty programs like BugCrowd or HackerOne deliver value: only telling you when there is something real. It can be a full time job to separate the wheat from the chaff. It’s made worse by the incentive of the reporters to make everything sound like a P1 hair-on-fire issue.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#84

Earlier quoted context omitted.

The users' comment history does read like generic LLM output. Look at the first lines of different comments: > Interesting point about Cranelift! I've been following its development for a while, and it seems like there's always something new popping up. > Interesting point about the color analysis! It kinda reminds me of how album art used to be such a significant part of music culture. > Interesting point about the…

> I suspect most of these cases aren't bots, they're users who put their thoughts, possibly in another language, into an LLM and ask it to form the comment for them. They like the text they see so they copy and paste it into HN. Yes, if this is LLM then it definitely wouldn't be zero-shot. I'm still on the fence myself as I've seen similar writing patterns with Asperger's (specifically what used to be called Asperger…

That's ye olde memetic "immune system" of the "onlygroup" (encapsulated ingroup kept unaware it's just an ingroup). "It don't sound like how we're taught, so we have no idea what it mean or why it there! Go back to Uncanny Valley!"

It's always enlightening to remember where Hans Asperger worked, and under what sociocultural circumstances that absolutely proverbial syndrome was first conceived.

GP evidently has some very subtle sort of expectations as to what authentic human expression must look like, which however seem to extend only as far as things like word choice and word order. (If that's all you ever notice about words, congrats, you're either a replicant or have a bad case of "learned literacy in USA" syndrome.)

This makes me want to point out that neither the means nor the purpose of the kind of communication which GP seems to implicitly expect (from random strangers) are even considered to be a real thing in many places and by many people.

I do happen to find that sort of thing way more coughinterestingcough than the whole "howdy stranger, are you AI or just a pseud" routine that HN posters seem to get such a huge kick out of.

Sure looks like one of the most basic moves of ideological manipulation: how about we solved the Turing Test "the wrong way around" by reducing the tester's ability to tell apart human from machine output, instead of building a more convincing language machine? Yay, expectations subverted! (While, in reality, both happen simultaneously.)

Disclaimer: this post was written by a certified paperclip optimizer.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#85
post #8
post #4

[flagged]

It's a little hilarious. First, as an organization, do all this cybersecurity theatre, and then create an MCP/LLM wormhole that bypasses it all. All because non-technical folks wave their hands about AI and not understanding the most fundamental reality about LLM software being fundamentally so different than all the software before it that it becomes an unavoidable black hole. I'm also a little pleased I used two sp…

Nitpick, but wormholes and black holes aren't limited to space! (unless you go with the Rick & Morty definition where "there's literally everything in space")

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#86

This might be off topic since we are in topic of AI tool and on HackerNews. I've been pondering a long time how does one build a startup company in domain they are not familiar with but ... Just have this urge to 'crave a pie' in this space. For the longest time, I had this dream of starting or building a 'AI Legal Tech Company' -- big issue is, I don't work in legal space at all. I did some cold reach on lawfirm rel…

I think if you have no domain expertise or unique insight it will be quite hard to find a real pain point to solve, deliver a winning solution, and have the ability to sell it.

Not impossible, but very hard. And starting a company is hard enough as it is.

So 9/10 times the answer will be to partner with someone who understands the space and pain point, preferably one who has lived it, or find an easier problem to solve.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#87
It's so great that they allowed him to publish a technical blog post. I once discovered a big vulnerability in a listed consumer tech company -- exposing users' private messages and also allowing to impersonate any user. The company didn't allow me to write a public blogpost.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#88

Earlier quoted context omitted.

It's the natural response. AI fans are routinely injecting themselves into every conversation here to somehow talk about AI ("I bet an AI tool would have found the issue faster") and AI is forcing itself onto every product. Comments dissing anything that sounds even remotely like AI is the logical response of someone who is fed up.

Every other headline and conversation having ai is super annoying. But also, its super annoying to sift through people saying "the word critical was used, this is obviously ai!". not to mention it really fucking sucks when you're the person who wrote something and people start chanting "ai slop! ai slop!". like, how am i going to prove is not AI? I can't wait until ai gets good enough that no one can tell the differe…

LLMs will never get good enough that no one can tell the difference, because the technology is fundamentally incapable of it, nor will it ever completely disappear, because the technology has real use cases that can be run at a massive profit.

Since LLMs are here to stay, what we actually need is for humans to get better at recognising LLM slop, and stop allowing our communication spaces to be rotted by slop articles and slop comments. It's weird that people find this concept objectional. It was historically a given that if a spambot posted a copy-pasted message, the comment would be flagged and removed. Now the spambot comments are randomly generated, and we're okay with it because it appears vaguely-but-not-actually-human-like. That conversations are devolving into this is actually the failure of HN moderation for allowing spambots to proliferate unscathed, rather than the users calling out the most blatantly obvious cases.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#89

It's so great that they allowed him to publish a technical blog post. I once discovered a big vulnerability in a listed consumer tech company -- exposing users' private messages and also allowing to impersonate any user. The company didn't allow me to write a public blogpost.

"Allow"?

Go on write your blog post. Don't let your dreams be dreams.

Re: Reverse engineering a $1B Legal AI tool exposed 100k+ confidential files

#90
post #77

Earlier quoted context omitted.

A lot of the time it’s less “nobody checked the security inbox” and more “the one person who understands that part of the system is juggling twelve other fires.” Security fixes are often a one-hour patch wrapped in two weeks of internal routing, approvals, and “who even owns this code?” archaeology. Holiday schedules and spam filters don’t help, but organizational entropy is usually the real culprit.

I've once had a whole sector of a fintech go down because one DevOps person ignored daily warning emails for three months that an API key was about to expire and needed reset. And of course nobody remembered the setup, and logging was only accessible by the same person, so figuring out also took weeks.

I'm currently on the other side of this trying to convince management that the maintenance that should have been done 3 years ago needs to get done. They need "justification".
Post reply on HN