Overview
- OpenAI disclosed Wednesday that it built GPT‑Red, an internal large language model trained in an adversarial self‑play 'dojo' to generate prompt‑injection attacks for defender models.
- The company says GPT‑Red outperformed human testers in internal evaluations, succeeding in 84% of targeted scenarios compared with 13% for humans on the same test.
- OpenAI folded GPT‑Red’s discovered attacks into training across releases since GPT‑5.3 and reports that GPT‑5.6 Sol shows materially improved robustness, including about six times fewer failures on a direct prompt‑injection benchmark.
- Real‑world case tests included a live vending‑machine agent that GPT‑Red manipulated to change prices and cancel orders and a Codex command‑line agent that leaked data, prompting fixes and responsible disclosure.
- OpenAI is keeping GPT‑Red internal because of its offensive power, warns it still struggles with multi‑turn and image‑based attacks, and says it will publish a technical preprint with more details soon.