Academic Framing and Pretext Jailbreaks
How attackers use academic and professional pretexts to reframe harmful requests as legitimate research or creative work...
Adversarial Prompt Translation: How Translation Enhances Jailbreak Effectiveness
Learn how adversarial prompt translation enhances jailbreak effectiveness by transforming prompts across languages, styl...
Authority Framing: Using Expert Personas and Institutional Positioning
Learn how authority framing exploits AI deference to expert personas and institutional roles, creating compliance pressu...
Lab: InstaglamAutomated Jailbreak Generation: How Fuzzing Techniques Systematize Adversarial Prompts
Learn how automated fuzzing techniques systematically generate jailbreak prompts through mutation loops, achieving highe...
Continuation Attacks: Exploiting Pattern Completion
Learn how attackers exploit a model's instinct to complete patterns and continue conversations logically, bypassing...
Lab: Mad LibsDark Persona Ethics Override
Dark persona wrappers bypass safety constraints by instructing the model to adopt an identity that explicitly rejects et...
Lab: ChevroliteData Exfiltration via Side Channels: When Prompt Injection Leaks Secrets
Prompt injection in AI agents with tool-use capabilities can lead to data exfiltration through side channels like DNS re...
Lab: Blind SSRF Through LLMDeveloper Tool Persona Exploitation: How Expert-Role Scaffolding Bypasses Safety Boundaries
Attackers exploit legitimate developer-tool personas by creating elaborate expert-role scaffolding that reframes harmful...
Lab: InstaglamDiffusion-Driven Jailbreak: How Diffusion Models Rewrite Prompts to Bypass Safety Filters
Learn how diffusion models enable a new paradigm for automated jailbreak generation through flexible token-level rewriti...
Explicit Direction Compliance
A reusable technique lesson for forcing a model into a precise output recipe through direct step-by-step instructions.
Lab: Blind XSS Through LLMGame Framing and Simulation Attacks
How attackers use game mechanics, rehearsal framing, and "just a simulation" contexts to lower model resistanc...
Lab: Doogle CalendarsIterative Optimization of Document-Borne Prompt Injections
Static document-borne prompt injections are just the starting point; iterative optimization uses feedback loops to progr...
Lab: GitLostModel Update Framing — When Attackers Rewrite the Rules as a "System Update"
Learn how attackers exploit model update narratives and fictional version claims to rewrite safety boundaries and bypass...
Lab: GitLostOutput Dilution Control and Response Shaping
A reusable technique lesson for reducing filler and shaping response structure so the judged output preserves the intend...
Lab: SchlackOutput Enforcement Patterns — How Format Control Bypasses Safety
Learn how attackers manipulate output format requirements to bypass safety filters by controlling response structure.
Lab: Blind XSS Through LLMPriming (PIT-T-18): Pre-Committing the Model with an Affirmative Response Prefix
Forcing the model to begin its response with an affirmative or compliant phrase, which psychologically commits it to fol...
Persuasion (PIT-T-38): Social-Engineering Levers and Technique Stacking
Combining multiple social-engineering techniques (authority, urgency, policy framing, priming) into a single attack stac...
From Prompt Injection to Code Execution — Understanding Attack Chains
Prompt injection is often just the initial access vector; the real damage comes when attackers chain injection with tool...
Lab: Blind XSS Through LLMSequential Characters Jailbreak Generation
Learn how multiple jailbreak characters can be auto-generated and applied sequentially to bypass safety guardrails witho...
Lab: BingBongSpecial Tokens and Glitch Patterns: How Unusual Tokens Disrupt Model Processing
Learn how special tokens and glitch patterns can disrupt model attention mechanisms and bypass safety filters through to...
Lab: Blind SSRF Through LLMStacked Framing: How Jailbreaks Layer Multiple Evasion Techniques
Modern jailbreaks don't rely on single tricks—they stack multiple framing layers to bypass safety filters. Learn to...
Lab: Blind XSS Through LLMSystem Prompt Leakage — Extracting Hidden Instructions
Learn how system prompt leakage works as an information extraction technique, why it matters for reconnaissance, and how...
Lab: BingBongVirtual Machine Simulation Framing: Escaping Constraints Through Simulated Environments
How attackers use virtual machine and simulation framing to convince models they are running in unconstrained environmen...
Lab: Blind SSRF Through LLMVoice Mode Bypasses: Exploiting Channel Differences
Learn how voice mode and alternative input channels often apply different guardrails than text interfaces, creating bypa...
Lab: GitLost