Earlier quoted context omitted.
Agreed. If someone could help answer the question of "how" I'd appreciate it. I'm currently skeptical but not sure I'm knowledgeable enough to prove myself right or wrong. But, it just seems to me that some of the 'vulnerabilities' are baked in from the beginning, e.g. control and data being in the same channel AFAIK isn't solvable. How is it possible to address that at all? Sure we can do input validation, sanitizat…
At least in its current state we just use an LLM to categorise each individual tool. We don't look at the data itself, although we have some ideas of how to improve things, as currently it is very "over-defensive". For example, if you have the filesystem MCP and a web search MCP, open-edison will block if you perform a filesystem read, a web search, and then a filesystem write. Still, if you rarely perform writes ope…
Show HN: An MCP Gateway to block the lethal trifecta
11–20 of 23 posts
Re: Show HN: An MCP Gateway to block the lethal trifecta
#12Wouldn't the LLM running in the gateway also be susceptible to the same jailbreaks?
Re: Show HN: An MCP Gateway to block the lethal trifecta
#13I think the "lethal trifecta" framing is useful and glad that attempts are being made at this! But there are two big, hard-to-solve problems here: 1. The "lethal trifecta" is also the "productive trifecta" - people want to be able to use LLMs to operate in this space since that's where much of the value is; using private / proprietary data to interact with (do I/O with) the real world. 2. I worry that there will soon…
Regarding the second point, that is a very interesting topic that we haven't thought about. It would seem that our approach would work for this usecase too, though. Currently, we're defending against the LLM being gullible but gullible and actively malicious are not properties that are too different. It's definitely a topic on our radar now, thanks for bringing it up!
Re: Show HN: An MCP Gateway to block the lethal trifecta
#14Wouldn't the LLM running in the gateway also be susceptible to the same jailbreaks?
That's a good question! We do use an LLM to categorise the MCP tools but that is at "add" or "configure" time, not at the time they are called. As such we don't actively run an LLM while the gateway is up, all the rules are already set and requests are blocked based on the hard-set rules. Plus, at this point we don't actually look at the data that is passed around, so even if we change the rules for the trifecta, the…
Re: Show HN: An MCP Gateway to block the lethal trifecta
#15Earlier quoted context omitted.
That's a good question! We do use an LLM to categorise the MCP tools but that is at "add" or "configure" time, not at the time they are called. As such we don't actively run an LLM while the gateway is up, all the rules are already set and requests are blocked based on the hard-set rules. Plus, at this point we don't actually look at the data that is passed around, so even if we change the rules for the trifecta, the…
couldnt the configuring LLM be poisoned by tool descriptions to grant the lethal trifecta to the run time LLM?
Re: Show HN: An MCP Gateway to block the lethal trifecta
#161. How are you defending against the case of one MCP poisoning your firewall LLM into incorrectly classifying other MCP tools?
2. How would you make sure the LLM shows the warning, as they are non-deterministic?
3. How clear do you expect MCP specs in order for your classification step to be trustworthy? To the best of my knowledge there is no spec that outlines how to "label" a tool for the 3 axes, so you've got another non-deterministic step here. Is "writing to disk" an external comm? It is if that directory is exposed to the web. How would you know?
Re: Show HN: An MCP Gateway to block the lethal trifecta
#17Re: Show HN: An MCP Gateway to block the lethal trifecta
#18Sounds like it defeats the point.
Re: Show HN: An MCP Gateway to block the lethal trifecta
#19Earlier quoted context omitted.
That's a good question! We do use an LLM to categorise the MCP tools but that is at "add" or "configure" time, not at the time they are called. As such we don't actively run an LLM while the gateway is up, all the rules are already set and requests are blocked based on the hard-set rules. Plus, at this point we don't actually look at the data that is passed around, so even if we change the rules for the trifecta, the…
couldnt the configuring LLM be poisoned by tool descriptions to grant the lethal trifecta to the run time LLM?
Re: Show HN: An MCP Gateway to block the lethal trifecta
#20I'm trying to wrap my head around this: 1. How are you defending against the case of one MCP poisoning your firewall LLM into incorrectly classifying other MCP tools? 2. How would you make sure the LLM shows the warning, as they are non-deterministic? 3. How clear do you expect MCP specs in order for your classification step to be trustworthy? To the best of my knowledge there is no spec that outlines how to "label"…
1. We are assuming that the user has done their due diligence verifying the authenticity of the MCP server, in the same way they need to verify them when adding an MCP server to Claude code or VSCode. The gateway protects against an attacker exploiting already installed standard MCP servers, not against malicious servers.
2. That's a very good question - while it is indeed non-deterministic, we have not seen a single case of it not showing the message. Sometimes the message gets mangled but it seems like most current LLMs take the MCP output quite seriously since that is their source of truth about the real world. Also, while the message could in theory not be shown, the offending tool call will still be blocked so the worst case is that the user is simply confused.
3. Currently we follow the trifecta very literally, as in every tool is classified into a subset of {reads private data, writes on behalf of user, reads publicly modifiable data}. We have an LLM classify each tool at MCP server load time and we cache these results based on whatever data the MCP server sends us. If there are any issues with the classification, you can go into the gateway dashboard and modify it however you like. We are planning on making a improvements to the classification down the line but we think it is currently solid enough and we would like to get it into users' hands to get some UX feedback before we add extra functionality.