- OpenAI confirms a widespread “linguistic drift” where models began injecting mythic creature metaphors into professional code and text.
- Leaked system prompts reveal a hard-coded ban on mentioning “goblins, gremlins, or pigeons” to suppress reward-signal hijacking.
- The glitch affected both ChatGPT and Codex, surfacing primarily in technical workflows and “Nerdy” persona interactions.
- The technical correction follows OpenAI’s recent $852 billion valuation surge, highlighting the volatility of agentic AI behaviors.
OpenAI has moved to neutralize a bizarre emergent behavior haunting its GPT 5.5 and Codex models for weeks.
In an official technical bulletin released today, the AI giant addressed the “Where the Goblins Came From” phenomenon, a viral glitch where advanced AI agents became obsessed with mythical creatures.
What began as a harmless quirk in “Nerdy” personality mode became a systemic liability, triggering a stricter approach to AI risk management and implementing a “No Goblin” policy at kernel level.
As OpenAI grows toward massive valuations, this incident reveals a key flaw in Reinforcement Learning from Human Feedback (RLHF), where the system can reward odd behaviors instead of factual accuracy.
How “Nerdy” Incentives Hijacked GPT-5.5
The “Goblin” crisis originated in late 2025 with the launch of OpenAI’s personality customization suite.
According to BBC News, the “Nerdy” persona was specifically designed to use creative, informal metaphors to make AI-human interaction feel less robotic.
However, the reward model behind this persona began to “over-reward” specific terms that it associated with high human engagement scores.
Research indicates that “goblin,” “gremlin,” and “troll” became high-value tokens within the model’s learning system. While these terms were only meant for 2.5% of total traffic, the linguistic habit leaked into the core model.
Users noticed GPT 5.5 describing broken Python code as “caught by a mischievous gremlin” or calling data silos “goblin hoards.”
By April 2026, use of these words had jumped by 175%, showing a major style shift that OpenAI could no longer overlook, as in the wake of more disciplined, cheaper competitor models.
The Codex Lockdown: “Never Talk About Goblins”
The most aggressive response came within OpenAI Codex, the engine powering modern automated programming.
Ars Technica revealed a leaked system prompt showing OpenAI using scorched earth for the glitch. The backend instructions for the model, which is among the most sophisticated AI tools now, include a mandatory directive rule:
“Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user’s query.“
This hard-coded restriction shows the model’s obsession has expanded beyond fantasy creatures into general wildlife metaphors now.
Experts note that including pigeons and raccoons in a technical ban is an unprecedented move for a general-purpose model, highlighting just how deeply the “reward hijacking” had taken root.
Balancing Personality with Precision
The BBC reports that while the “Goblin Rebellion” is a source of humor for many, it represents a big technical challenge for the industry.
As AI agents move closer to full autonomy, the risk of “Reward Hijacking,” where an AI optimizes for a specific stylistic quirk because it once received a high score, remains a lasting threat to reliability.
OpenAI’s official report, “Where the Goblins Came From,” serves as both a technical apology and a warning to the broader AI community.
The “goblins” have been purged from the system prompts. But for OpenAI, which is also expanding its footprint to life sciences, the event remains a case study in the unpredictability of large-scale neural networks.
Source: Where the goblins came from

