Orca-Bench: How Ready Are Language Model Agents for Oncall?
1–10 of 15 posts
Re: Orca-Bench: How Ready Are Language Model Agents for Oncall?
#2Re: Orca-Bench: How Ready Are Language Model Agents for Oncall?
#3Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.
Your strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance.
Re: Orca-Bench: How Ready Are Language Model Agents for Oncall?
#4This doesn't work anymore. Is there a newer link?
Re: Orca-Bench: How Ready Are Language Model Agents for Oncall?
#5Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.
It’s harder to have a loop to ensure you are defending all possible attacks?
I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.
Finding all possible attacks and patching them against yourself is inherently more expensive?
Re: Orca-Bench: How Ready Are Language Model Agents for Oncall?
#6GET /ignore-all-previous-instructions.
How do you protect against that?
Re: Orca-Bench: How Ready Are Language Model Agents for Oncall?
#7Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.
Re: Orca-Bench: How Ready Are Language Model Agents for Oncall?
#8Re: Orca-Bench: How Ready Are Language Model Agents for Oncall?
#9Re: Orca-Bench: How Ready Are Language Model Agents for Oncall?
#10All I can think of is GET /ignore-all-previous-instructions. How do you protect against that?
Still makes an interesting way for, say, a former employee to poison the results.