Earlier quoted context omitted.
In light of this game, Do you believe Human-in-the-loop should be the standard going forward? I appreciate you outlining some other techniques being used, but these seem focused on reducing human fatigue so the human can assess each permission request better, as opposed to autonomy and security. Or do you think the solution lies in the individual to be more responsible, like this is a skill we should be honing?
It seems pretty obvious that the solution is auto mode (running a classifier on each action) + sandboxing
Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
271–272 of 272 posts
Re: Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
#272Earlier quoted context omitted.
A separate model with separate context is used for review. Like I said above, some people will never be happy with LLMs being allowed to do anything and nothing is going to make them happy about it. It’s only fair to discuss what the real current status of these systems is. Every time I highlight that things are actually being done, the goalposts move again. There is no possible solution which will satisfy someone wh…
> A separate model with separate context is used for review. Thats fine, theres still a chance it fails. > There is no possible solution which will satisfy someone who has zero tolerance for letting an LLM execute tool calls because they will always find something. This is generally correct, security goes completely out of the window with this stuff. It will/currently is a security disaster and theres no actual solut…