Overview
- OpenAI disclosed July 15 that it built GPT-Red, an internal large language model trained in adversarial self-play to generate and refine prompt-injection attacks against defender models.
- GPT-Red ran attacker-defender loops in a simulated 'dojo' that let it probe browsing, code editing, and agent workflows and surface new attack patterns for developers to fix.
- One new method GPT-Red exposed is a 'fake chain of thought,' which inserts spoofed internal reasoning into a model’s prompts to trick it into following false instructions.
- OpenAI says attacks GPT-Red produced worked against earlier models at high rates but that feeding those attacks into training cut GPT-5.6’s prompt-injection failures sharply according to the company’s internal benchmarks.
- The company will keep GPT-Red private, stresses the tool complements human and third-party red teams, and acknowledges limits such as weaker performance on multi-turn conversational and image-based attacks while independent verification of OpenAI’s metrics is still pending.