GPT‑Red: Unlocking Self-Improvement for Robustness
1–2 of 2 posts
Re: GPT‑Red: Unlocking Self-Improvement for Robustness
#2Useful direction, but the hard part seems to be measuring novelty after each fix. Are they reporting whether later red-team cases are genuinely distinct, or mostly variants of the same failure mode?