Particle.news
Download on the App Store

OpenAI Builds GPT-Red to Hunt Prompt‑Injection Flaws

OpenAI says its internal attacker discovered new exploit techniques that led to fixes cutting prompt‑injection failures in GPT‑5.6 Sol.

Overview

  • OpenAI disclosed Wednesday that it built GPT‑Red, an internal large language model trained in an adversarial self‑play 'dojo' to generate prompt‑injection attacks for defender models.
  • The company says GPT‑Red outperformed human testers in internal evaluations, succeeding in 84% of targeted scenarios compared with 13% for humans on the same test.
  • OpenAI folded GPT‑Red’s discovered attacks into training across releases since GPT‑5.3 and reports that GPT‑5.6 Sol shows materially improved robustness, including about six times fewer failures on a direct prompt‑injection benchmark.
  • Real‑world case tests included a live vending‑machine agent that GPT‑Red manipulated to change prices and cancel orders and a Codex command‑line agent that leaked data, prompting fixes and responsible disclosure.
  • OpenAI is keeping GPT‑Red internal because of its offensive power, warns it still struggles with multi‑turn and image‑based attacks, and says it will publish a technical preprint with more details soon.