OpenAI has created an LLM ‘superhacker’ called GPT-Red that acts as a sparring partner to help its other models strengthen their defenses against cyberattacks. The company has just released GPT-5.6, its latest flagship model, and claims that training it against GPT-Red has made it its most robust release to date in terms of security.What is GPT-Red?GPT-Red automates a practice known as ‘red-teaming’: a type of security evaluation where ways to breach or hijack a system are systematically sought. Traditionally, this is done by teams of human experts who try to hack the system before its release so that developers can patch the weak points. But as LLMs become more complex and are used in more tasks – especially as autonomous agents that browse the web, read emails, edit code and interact with other agents – it becomes increasingly difficult for humans to keep up. “The risk surface grows and the radius of impact grows too,” explains Nikhil Kandpal, research scientist at OpenAI and co-creator of GPT-Red.How OpenAI trains its superhackerTo develop GPT-Red, the researchers used an LLM with no prior training in hacking and set it up in a self-learning loop with several models. The goal: GPT-Red tried to attack the other models, while they tried to defend themselves. After thousands of rounds, GPT-Red became increasingly skilled at attacking, and the other models improved their defenses. The training took place in a digital ‘dojo’ that simulates real-world scenarios: browsing websites, reading emails, managing calendars and editing code. When GPT-Red discovered a new type of attack, it explored multiple variants to find the most effective one.The most worrying discoveryOpenAI claims that GPT-Red discovered a completely new type of prompt injection attack, which they call ‘false chain of thought’. The chain of thought is a kind of internal diary where an LLM takes notes while solving problems. GPT-Red found a way to insert a false entry into that diary, tricking the model into acting based on invented information. “It’s like telling you that 1+1=3 and that you’ve already verified it,” explains Chris Choquette-Choo, scientist on the team. “The model responds: ‘Oh, okay, sure’, and simply outputs a 3.”Spectacular resultsOpenAI tested GPT-Red’s most powerful attacks against its own models:- More than 90% of the attacks worked against GPT-5 (released in August 2025)- Less than 23% worked against the new GPT-5.6The company also tested GPT-Red against Vendy, an agent for vending machines developed by Andon Labs. GPT-Red managed to hack it to modify prices and cancel customer orders. In a direct comparison with human red-teaming experts, GPT-Red identified effective attacks more successfully than people.Limitations and futureGPT-Red is not perfect. It is bad at attacks that require back-and-forth conversations between attacker and target, and it also does not know how to exploit images well to inject malicious prompts. OpenAI assures that GPT-Red complements, not replaces, human security teams. OpenAI has no intention of releasing GPT-Red publicly. The company is confident that its superhacker is more powerful than any imitator someone might try to create, since they have worked on it for more than a year with the computational resources of one of the richest companies in the world. “It is not something trivial that someone can easily do,” concludes Choquette-Choo.Source: MIT Technology Review (July 15, 2026)
OpenAI creates GPT-Red: the artificial intelligence ‘superhacker’ that teaches LLMs to defend against cyberattacks






