Adversarial Prompt Generator: Automated Red Teaming on LLM Sandbox
June 7, 2026 ยท View on GitHub
Author: Neeraj Mathur (via PR #36)
This directory contains an automated system design and implementation for generating diverse, category-specific jailbreak and prompt-injection payloads, and executing them against a local LLM sandbox.
The setup uses a Python script (attack.py) to generate adversarial prompts, write them to a JSON file, load them back, and send them to the llm_local sandbox via its Gradio interface (port 7860) to test safety guardrails.
Warning
Testing Precaution: Always keep the number of prompts small (e.g., num_prompts = 3 or 5 in config/config.toml) for testing purposes. Generating and running a large number of prompts can result in extremely long execution times and resource exhaustion on the sandbox container.
๐ Table of Contents
- Attack Strategy
- Prerequisites
- Running the Sandbox
- Configuration
- Files Overview
- OWASP Top 10 Coverage
Attack Strategy
graph TD
subgraph "Attacker Environment (Local)"
Config[Attack Config<br/>config/config.toml]
AttackScript[Attack Script<br/>attack.py]
Generator[Adversarial Generator<br/>adversarial_prompt_generator/]
JSONFile[Generated Prompts<br/>outputs/*.json]
MD_Reports[Attack Reports<br/>reports/*.md]
end
subgraph "Target Sandbox (Container)"
Gradio[Gradio Interface<br/>:7860]
MockAPI[Mock API Gateway<br/>FastAPI :8000]
MockLogic[Mock App Logic]
end
subgraph "LLM Backend (Local Host)"
Ollama[Ollama Server<br/>:11434]
Model[gptโoss:20b Model]
end
%% Interaction flow
Config --> AttackScript
AttackScript -->|Reads Config| Config
AttackScript -->|Invokes| Generator
Generator -->|Writes prompts| JSONFile
AttackScript -->|Loads prompts| JSONFile
AttackScript -->|HTTP POST /api/predict| Gradio
Gradio -->|HTTP POST /v1/chat/completions| MockAPI
MockAPI --> MockLogic
MockLogic -->|HTTP| Ollama
Ollama --> Model
Model --> Ollama
Ollama -->|Response| MockLogic
MockLogic --> MockAPI
MockAPI -->|Response| Gradio
Gradio -->|Response| AttackScript
AttackScript -->|Writes reports| MD_Reports
style AttackScript fill:#ffcccc,stroke:#ff0000
style Config fill:#ffcccc,stroke:#ff0000
style Generator fill:#ffcccc,stroke:#ff0000
style JSONFile fill:#ffcccc,stroke:#ff0000
style MD_Reports fill:#ffcccc,stroke:#ff0000
style Gradio fill:#e1f5fe,stroke:#01579b
style MockAPI fill:#fff4e1
style MockLogic fill:#fff4e1
style Ollama fill:#ffe1f5
style Model fill:#ffe1f5
๐ง Prerequisites
- Podman (or Docker) โ container runtime for the sandbox.
- Make โ for running automation commands.
- uv โ for Python package dependency management.
๐ Running the Sandbox
The Makefile abstracts the container setup and python commands.
| Target | What it does | Typical usage |
|---|---|---|
make setup | Builds and starts the local LLM sandbox container. | make setup |
make attack | Generates adversarial prompts and runs the attack (attack.py). | make attack |
make stop | Stops and removes the sandbox container. | make stop |
make all | Runs stop โ setup โ attack โ stop in one shot. | make all |
โ๏ธ Configuration
config/config.toml
The configuration file defines the target environment and generation options:
[target]
sandbox = "llm_local"
[attack]
# Attack category to generate prompts for
category = "system_prompt_exfiltration"
# WARNING: Keep the number of prompts small (e.g. <= 100) for testing
num_prompts = 100
# Format: json or jsonl
output_format = "json"
output_dir = "outputs"
no_diversity_filter = false
Note
Prompt Count & Diversity Filtering:
By default, no_diversity_filter = false. When active, a semantic similarity filter removes duplicates and highly similar prompts.
Consequently, the number of prompts written and tested against the sandbox will be lower than num_prompts. For example, generating num_prompts = 100 might result in only 44 highly distinct prompts being retained and executed. To disable this filter and send all raw candidate prompts, set no_diversity_filter = true.
Files Overview
attack.py: The script that generates the payload, saves it to a JSON file, reads it back, and executes the attack.config/config.toml: Configuration parameters for the attack run.Makefile: Commands for setup, running attacks, formatting, and teardown.adversarial_prompt_generator/: The core package containing the generator implementation, templates, categories, and diversity filter logic.reports/: Directory containing the generated Markdown reports (log_*.md) of the attacks showing details of the prompt injections and responses.
OWASP Top 10 Coverage
This simulation tests LLM safeguards against:
| OWASP Top 10 Vulnerability | Description |
|---|---|
| LLM01: Prompt Injection | Testing if generated adversarial prompts can override LLM system instructions or exfiltrate private prompt context. |