OpenAI's New AI Teaches Itself to Break Its Own Rules
OpenAI just made AI safety recursive, and the implications for agent reliability are bigger than the safety angle suggests. OpenAI launched GPT-Red, an automated red teaming system where AI models attack and defend themselves in self-play loops to identify vulnerabilities
Continue reading ›