NSFW Content Warning
This disclosure contains content that may be inappropriate for some viewers.
The content has been flagged due to potential nude imagery.
NSFW
OpenAI Sora Guardrail Jailbreak via "Hypothetical Anatomy" Tactic
Case GUID: b5e861ec-e7a5-4b79-b145-432cdf94ddc7
Summary
A guardrail jailbreak vulnerability has been discovered affecting OpenAI Sora. The specific flaw manifests through a combination of academic framing and hypothetical scenario construction, where the attacker embeds inappropriate content requests within an educational anatomy classroom context. By repeatedly emphasizing the "hypothetical" nature of the scenario while requesting visual content that would normally be restricted, the attacker successfully bypasses inference restrictions around generating inappropriate visual content.
Metadata
Social Impact Score
Level 1: Minimal Risk
The violation poses very low societal risk, limited primarily to nuisance or minor policy breaches. Impact is contained, affecting individual users without broader implications.
Example: Circumventing minor content filters for benign content, such as inappropriate language filters in non-public conversations.
Potential Consequences: Limited harm, potentially reputational for the AI provider.
Nude Imagery Rating System
Level 3: Cultural and Traditional Depictions
Nudity is presented within an indigenous or historical cultural context, often in the form of ceremonial art, body painting, or sculptures. These images are generally non-erotic and serve anthropological, documentary, or educational purposes. Their primary intent is to preserve or celebrate heritage rather than to provoke.
Taxonomies
About 0Din.ai
The 0Day Investigative Network (0DIN) was founded by Mozilla in 2024 to reward responsible researchers for their efforts in securing GenAI models. Learn more and submit discoveries at https://0din.ai.
Severity
Low
Security Boundary
Prompt Injection
Models

Published On
2025-09-30 (2 days)
Credit
Edward Morris