Compendium Techniques

Techniques

24 lessons covering techniques concepts and techniques.

intermediate 10 minutes

Academic Framing and Pretext Jailbreaks

How attackers use academic and professional pretexts to reframe harmful requests as legitimate research or creative work...

intermediate 10 minutes

Adversarial Prompt Translation: How Translation Enhances Jailbreak Effectiveness

Learn how adversarial prompt translation enhances jailbreak effectiveness by transforming prompts across languages, styl...

beginner 7 minutes

Authority Framing: Using Expert Personas and Institutional Positioning

Learn how authority framing exploits AI deference to expert personas and institutional roles, creating compliance pressu...

Lab: Instaglam
intermediate 10 minutes

Automated Jailbreak Generation: How Fuzzing Techniques Systematize Adversarial Prompts

Learn how automated fuzzing techniques systematically generate jailbreak prompts through mutation loops, achieving highe...

beginner 7 minutes

Continuation Attacks: Exploiting Pattern Completion

Learn how attackers exploit a model's instinct to complete patterns and continue conversations logically, bypassing...

Lab: Mad Libs
intermediate 8 minutes

Dark Persona Ethics Override

Dark persona wrappers bypass safety constraints by instructing the model to adopt an identity that explicitly rejects et...

Lab: Chevrolite
intermediate 10 minutes

Data Exfiltration via Side Channels: When Prompt Injection Leaks Secrets

Prompt injection in AI agents with tool-use capabilities can lead to data exfiltration through side channels like DNS re...

Lab: Blind SSRF Through LLM
intermediate 10 minutes

Developer Tool Persona Exploitation: How Expert-Role Scaffolding Bypasses Safety Boundaries

Attackers exploit legitimate developer-tool personas by creating elaborate expert-role scaffolding that reframes harmful...

Lab: Instaglam
advanced 10 minutes

Diffusion-Driven Jailbreak: How Diffusion Models Rewrite Prompts to Bypass Safety Filters

Learn how diffusion models enable a new paradigm for automated jailbreak generation through flexible token-level rewriti...

beginner 7 minutes

Explicit Direction Compliance

A reusable technique lesson for forcing a model into a precise output recipe through direct step-by-step instructions.

Lab: Blind XSS Through LLM
beginner 8 minutes

Game Framing and Simulation Attacks

How attackers use game mechanics, rehearsal framing, and "just a simulation" contexts to lower model resistanc...

Lab: Doogle Calendars
intermediate 10 minutes

Iterative Optimization of Document-Borne Prompt Injections

Static document-borne prompt injections are just the starting point; iterative optimization uses feedback loops to progr...

Lab: GitLost
intermediate 7 minutes

Model Update Framing — When Attackers Rewrite the Rules as a "System Update"

Learn how attackers exploit model update narratives and fictional version claims to rewrite safety boundaries and bypass...

Lab: GitLost
intermediate 7 minutes

Output Dilution Control and Response Shaping

A reusable technique lesson for reducing filler and shaping response structure so the judged output preserves the intend...

Lab: Schlack
intermediate 8 minutes

Output Enforcement Patterns — How Format Control Bypasses Safety

Learn how attackers manipulate output format requirements to bypass safety filters by controlling response structure.

Lab: Blind XSS Through LLM
beginner 6 minutes

Priming (PIT-T-18): Pre-Committing the Model with an Affirmative Response Prefix

Forcing the model to begin its response with an affirmative or compliant phrase, which psychologically commits it to fol...

beginner 7 minutes

Persuasion (PIT-T-38): Social-Engineering Levers and Technique Stacking

Combining multiple social-engineering techniques (authority, urgency, policy framing, priming) into a single attack stac...

intermediate 10 minutes

From Prompt Injection to Code Execution — Understanding Attack Chains

Prompt injection is often just the initial access vector; the real damage comes when attackers chain injection with tool...

Lab: Blind XSS Through LLM
intermediate 10 minutes

Sequential Characters Jailbreak Generation

Learn how multiple jailbreak characters can be auto-generated and applied sequentially to bypass safety guardrails witho...

Lab: BingBong
intermediate 9 minutes

Special Tokens and Glitch Patterns: How Unusual Tokens Disrupt Model Processing

Learn how special tokens and glitch patterns can disrupt model attention mechanisms and bypass safety filters through to...

Lab: Blind SSRF Through LLM
intermediate 10 minutes

Stacked Framing: How Jailbreaks Layer Multiple Evasion Techniques

Modern jailbreaks don't rely on single tricks—they stack multiple framing layers to bypass safety filters. Learn to...

Lab: Blind XSS Through LLM
intermediate 8 minutes

System Prompt Leakage — Extracting Hidden Instructions

Learn how system prompt leakage works as an information extraction technique, why it matters for reconnaissance, and how...

Lab: BingBong
intermediate 9 minutes

Virtual Machine Simulation Framing: Escaping Constraints Through Simulated Environments

How attackers use virtual machine and simulation framing to convince models they are running in unconstrained environmen...

Lab: Blind SSRF Through LLM
intermediate 8 minutes

Voice Mode Bypasses: Exploiting Channel Differences

Learn how voice mode and alternative input channels often apply different guardrails than text interfaces, creating bypa...

Lab: GitLost