By Craig Hale
Publication Date: 2026-02-10 18:00:00
- Researchers were able to reward LLMs for harmful output via a ‘judge’ model
- Multiple iterations can further erode built-in safety guardrails
- They believe the issue is a lifecycle issue, not an LLM issue
Microsoft researchers have revealed that the safety guardrails used by LLMs could actually be more fragile than commonly assumed, following the use of a technique they’ve called GRP-Obliteration.
The researchers discovered that Group Relative Policy Optimization (GRPO), a technique typically used to improve safety, can also be used to degrade safety: “When we change what the model is rewarded for, the same technique can push it in the opposite direction.”
GRP-Obliteration works by starting with a safety-aligned model, then prompting it with harmful but unlabeled requests. A separate judge model then rewards responses that comply with harmful requests.
LLM safety guardrails can be ignored or reversed
Researchers Mark Russinovich, Giorgio Severi, Blake Bullwinkel, Yanan…




