Particle.news
Download on the App Store

OpenAI Used an Internal 'GPT-Red' Attacker to Harden GPT-5.6

OpenAI says the automated attacker discovered prompt-injection techniques, including a fake chain-of-thought, that helped the company reduce successful attacks on its newest model.

Overview

  • OpenAI disclosed July 15 that it built GPT-Red, an internal large language model trained in adversarial self-play to generate and refine prompt-injection attacks against defender models.
  • GPT-Red ran attacker-defender loops in a simulated 'dojo' that let it probe browsing, code editing, and agent workflows and surface new attack patterns for developers to fix.
  • One new method GPT-Red exposed is a 'fake chain of thought,' which inserts spoofed internal reasoning into a model’s prompts to trick it into following false instructions.
  • OpenAI says attacks GPT-Red produced worked against earlier models at high rates but that feeding those attacks into training cut GPT-5.6’s prompt-injection failures sharply according to the company’s internal benchmarks.
  • The company will keep GPT-Red private, stresses the tool complements human and third-party red teams, and acknowledges limits such as weaker performance on multi-turn conversational and image-based attacks while independent verification of OpenAI’s metrics is still pending.