Image Generation Guide
September 16, 2026 · View on GitHub
Complete guide to generating and editing images using OpenAI's APIs.
Overview
Generate high-quality images for use with Sora video generation or standalone use. Two APIs are available:
| API | Tool | Best For |
|---|---|---|
| Images API | generate_image, edit_image | New generation with gpt-image-2.5 (RECOMMENDED) |
| Responses API | create_image | Iterative refinement with previous_response_id |
Choose Images API when:
- Creating new images from scratch
- Editing existing images
- You want gpt-image-2.5 (state-of-the-art quality, up to 4K output, transparent backgrounds)
- You want synchronous results (no polling)
Choose Responses API when:
- Building iterative refinement chains with
previous_response_id - Working with a mainline model's conversational image generation
- You want to leverage the
actionfield to force generate vs. edit
Images API with gpt-image-2.5 (Recommended)
The Images API provides synchronous image generation with OpenAI's newest gpt-image-2.5 models. This is the recommended path for new image generation.
gpt-image-2.5 ships as two variants, both priced like gpt-image-2:
| Variant | Default for | OpenAI's positioning |
|---|---|---|
gpt-image-2.5-flare | generate_image, create_image | Fast, high-quality everyday generation |
gpt-image-2.5-sunburst | edit_image | Workflows where editing precision matters most |
Both accept background="transparent" (with png or webp output) and two quality levels above high: "xhigh" and "max".
Key Advantages
- Synchronous: Returns immediately (no polling required)
- gpt-image-2.5: State-of-the-art quality, up to 4K output, accepts thousands of valid resolutions (same size rules as gpt-image-2)
- Token usage tracking: Monitor costs with detailed token counts
- Faster: ~3s single-pass generation vs. multi-pass legacy models
When to pick another model
- gpt-image-2: the previous flagship (~99% text accuracy). It does not support
background="transparent"orquality="xhigh"|"max", and it ignoresinput_fidelity(always high) — the wrappers raise or strip accordingly. - gpt-image-1.5: older, fixed sizes; still supports transparent output and is the only current model that honours
input_fidelity(gpt-image-2.5 rejects the flag like gpt-image-2).
Basic Generation
# Generate an image (returns immediately)
result = generate_image(prompt="sunset over mountains")
# File immediately available at result.filename
# With quality and size options
result = generate_image(
prompt="professional product photo, studio lighting",
size="1536x1024",
quality="high",
filename="product.png"
)
Image Editing
# Edit a single image
result = edit_image(
prompt="add a hat to the person",
input_images=["portrait.png"]
)
# Multi-image composition (up to 16 images)
result = edit_image(
prompt="create a gift basket containing all these items",
input_images=["lotion.png", "soap.png", "candle.png"]
)
# Masked inpainting
result = edit_image(
prompt="add a flamingo standing in the water",
input_images=["pool.png"],
mask_filename="pool_mask.png" # PNG with alpha channel
)
Token Usage Tracking
result = generate_image(prompt="cityscape at night")
if result.usage:
print(f"Input tokens: {result.usage.input_tokens}")
print(f"Output tokens: {result.usage.output_tokens}")
print(f"Total tokens: {result.usage.total_tokens}")
Complete Workflow: Generate → Animate
# 1. Generate reference image (synchronous, no polling!)
result = generate_image(
prompt="A lone astronaut standing on a red desert planet, cinematic lighting",
size="1536x1024",
quality="high",
filename="astronaut.png"
)
# 2. Prepare for Sora (resize if needed)
prep = prepare_reference_image(
"astronaut.png",
"1280x720",
resize_mode="crop"
)
# 3. Generate video
video = create_video(
prompt="The astronaut turns and walks toward the horizon",
size="1280x720",
input_reference_filename=prep.output_filename
)
Responses API with a mainline model
Use the Responses API when you need iterative refinement with previous_response_id. This creates a conversational workflow where each image builds on the previous.
Tip: the image model defaults to gpt-image-2.5-flare; pin another with tool_config={"type": "image_generation", "model": "gpt-image-2.5-sunburst"} (or "gpt-image-2"). Transparent backgrounds work on the default.
Basic Workflow
# 1. Generate image
resp = create_image(prompt="sunset over mountains")
# 2. Poll for completion
status = get_image_status(resp.id)
# Check until status.status == "completed"
# 3. Download to IMAGE_PATH
download_image(resp.id, filename="sunset.png")
# 4. Use with video generation
video = create_video(
prompt="The sun rises slowly over the peaks",
input_reference_filename="sunset.png",
size="1280x720"
)
Models
gpt-6-astra (Default)
create_image(prompt="...", model="gpt-6-astra") # Default
Best for:
- Iterative refinement workflows
- Conversational image generation
- Complex multi-step prompts
Characteristics:
- OpenAI's latest model
- Excellent prompt following
- Supports
previous_response_idchains
gpt-5.6-sol / gpt-5.6-terra / gpt-5.6-luna
create_image(prompt="...", model="gpt-5.6-terra") # balanced cost
create_image(prompt="...", model="gpt-5.6-luna") # cheapest
create_image(prompt="...", model="gpt-4.1")
Best for:
- Alternative interpretations
- Experimentation with different styles
Iterative Refinement
Use previous_response_id to refine images conversationally without starting over:
# 1. Generate initial concept
resp1 = create_image(prompt="a futuristic cityscape")
get_image_status(resp1.id) # wait for completion
# 2. Refine iteratively
resp2 = create_image(
prompt="add more neon lights and flying cars",
previous_response_id=resp1.id
)
get_image_status(resp2.id) # wait
# 3. Further refinement
resp3 = create_image(
prompt="make it nighttime with rain",
previous_response_id=resp2.id
)
get_image_status(resp3.id) # wait
# 4. Download final version
download_image(resp3.id, filename="cityscape_final.png")
Benefits of iterative refinement:
- Faster than regenerating from scratch
- Maintains consistent style and composition
- Fine-tune specific elements
- Conversational workflow
Parameters
Basic Parameters
create_image(
prompt="your description here", # Required
model="gpt-6-astra", # Optional: gpt-6-astra (default), gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna
previous_response_id="resp_123" # Optional: for refinement
)
Advanced Configuration
For advanced control, use tool_config parameter:
from openai.types.responses.tool_param import ImageGeneration
config = ImageGeneration(
type="image_generation",
model="gpt-image-2.5-flare", # Image model (default)
size="1536x1024", # Image dimensions
quality="high", # Quality level
output_format="png", # Format
moderation="auto", # Content moderation
action="auto", # auto | generate | edit
)
resp = create_image(
prompt="logo for tech startup",
model="gpt-5.6-terra", # mainline model driving the tool
tool_config=config
)
tool_config fields:
- model:
"gpt-image-2.5-flare"/"gpt-image-2.5-sunburst"(defaults; transparent + xhigh/max),"gpt-image-2"(previous flagship),"gpt-image-1.5"(supports transparent),"gpt-image-1","gpt-image-1-mini" - size:
"auto","1024x1024","1024x1536","1536x1024", plus"2048x2048"/"3840x2160"and any valid resolution on gpt-image-2 - quality:
"auto","low","medium","high" - output_format:
"png","jpeg","webp" - background:
"auto","transparent"(gpt-image-1/1.5 only),"opaque" - input_fidelity:
"high","low"(gpt-image-1/1.5 only — gpt-image-2 and 2.5 are always high) - moderation:
"auto"(default),"low"(more permissive) - action:
"auto"(default),"generate"(force new image),"edit"(force edit of in-context image) - partial_images:
0-3— stream partial images during generation
Image Editing
Edit existing images by providing input images:
Single Image Editing
# Place base image in IMAGE_PATH first
resp = create_image(
prompt="add a flamingo to the pool",
input_images=["pool.png"]
)
The tool will:
- Read
pool.pngfromIMAGE_PATH - Apply your edits (add flamingo)
- Generate new image
Multiple Image Composition
# Combine multiple images
resp = create_image(
prompt="create a gift basket with all these items",
input_images=["lotion.png", "soap.png", "candle.png"]
)
Image order matters:
- First image receives highest fidelity
- Use most important image first
- Up to multiple images supported
Masked Inpainting
# Edit specific region using mask
resp = create_image(
prompt="add logo to woman's shirt",
input_images=["woman.jpg", "logo.png"],
mask_filename="shirt_mask.png"
)
Mask requirements:
- PNG format with alpha channel
- Transparent = edit this area
- Black = keep original
- Mask must match first input image dimensions
Complete Workflows
Workflow 1: Generate Reference for Video
# Step 1: Generate base image
print("Generating reference image...")
resp = create_image(
prompt="A lone astronaut standing on a red desert planet, cinematic lighting",
model="gpt-5.6-terra"
)
# Step 2: Wait for completion
import time
while True:
status = get_image_status(resp.id)
if status.status == "completed":
break
elif status.status == "failed":
raise Exception("Image generation failed")
print(f"Status: {status.status}")
time.sleep(2)
# Step 3: Download
result = download_image(resp.id, filename="astronaut.png")
print(f"Downloaded: {result.filename}")
# Step 4: Prepare for video (if dimensions don't match)
prep = prepare_reference_image(
"astronaut.png",
"1280x720",
resize_mode="crop"
)
# Step 5: Generate video
video = create_video(
prompt="The astronaut turns and walks toward the horizon",
size="1280x720",
input_reference_filename=prep.output_filename
)
Workflow 2: Iterative Design Process
# Start with concept
print("Phase 1: Initial concept")
resp1 = create_image(prompt="modern minimalist logo for AI company")
get_image_status(resp1.id) # wait
# Add details
print("Phase 2: Adding details")
resp2 = create_image(
prompt="add blue and silver color scheme",
previous_response_id=resp1.id
)
get_image_status(resp2.id) # wait
# Refine style
print("Phase 3: Style refinement")
resp3 = create_image(
prompt="make it more geometric and abstract",
previous_response_id=resp2.id
)
get_image_status(resp3.id) # wait
# Final touches
print("Phase 4: Final touches")
resp4 = create_image(
prompt="add subtle gradient and increase contrast",
previous_response_id=resp3.id
)
get_image_status(resp4.id) # wait
# Download final
download_image(resp4.id, filename="logo_final.png")
Workflow 3: Product Visualization
# Generate base product image
resp = create_image(
prompt="luxury perfume bottle on marble surface, studio lighting",
model="gpt-5.6-terra"
)
get_image_status(resp.id) # wait
download_image(resp.id, filename="perfume_base.png")
# Create variant 1: Different angle
resp2 = create_image(
prompt="rotate 45 degrees, show side profile",
previous_response_id=resp.id
)
get_image_status(resp2.id) # wait
download_image(resp2.id, filename="perfume_side.png")
# Create variant 2: Different setting
resp3 = create_image(
prompt="place on dark velvet fabric with soft rim lighting",
previous_response_id=resp.id
)
get_image_status(resp3.id) # wait
download_image(resp3.id, filename="perfume_luxury.png")
Workflow 4: Logo Placement on Images
# You have: company_shirt.jpg, company_logo.png in IMAGE_PATH
# Add logo to shirt
resp = create_image(
prompt="place the logo on the chest area of the shirt",
input_images=["company_shirt.jpg", "company_logo.png"],
tool_config=ImageGeneration(
type="image_generation",
input_fidelity="high" # Preserve quality
)
)
get_image_status(resp.id) # wait
download_image(resp.id, filename="shirt_with_logo.png")
# Now animate it
video = create_video(
prompt="The person wearing the shirt turns and smiles at the camera",
input_reference_filename="shirt_with_logo.png",
size="1280x720"
)
Best Practices
Prompting Tips
Be specific:
# ❌ Vague
create_image(prompt="a nice landscape")
# ✅ Specific
create_image(prompt="mountain landscape at golden hour, pine trees in foreground, snow-capped peaks, volumetric fog")
Use artistic references:
create_image(prompt="portrait in the style of Rembrandt, dramatic lighting, oil painting texture")
Specify composition:
create_image(prompt="wide shot of city skyline, rule of thirds composition, sunset backlighting")
Quality Optimization
For highest quality:
config = ImageGeneration(
type="image_generation",
model="gpt-image-2.5-flare", # default
quality="high",
output_format="png" # Lossless
)
create_image(prompt="...", tool_config=config)
For speed:
config = ImageGeneration(
type="image_generation",
model="gpt-image-1-mini", # Faster model
quality="medium"
)
create_image(prompt="...", tool_config=config)
File Management
Use descriptive filenames:
download_image(resp.id, filename="project_hero_v3.png")
# Better than: image_123.png
Organize by project:
IMAGE_PATH/
├── project_alpha/
│ ├── hero_image.png
│ ├── logo_variants/
│ └── backgrounds/
└── project_beta/
└── references/
Error Handling
resp = create_image(prompt="...")
# Check status
status = get_image_status(resp.id)
if status.status == "failed":
print(f"Generation failed for: {resp.id}")
# Try with different prompt or settings
elif status.status == "completed":
download_image(resp.id)
else:
print(f"Still generating: {status.status}")
Common Issues
Image Not Completing
Problem: Status stuck at "in_progress" Solutions:
- Wait longer (complex images take time)
- Check OpenAI status page for service issues
- Try simpler prompt
- Use
generate_imageinstead (synchronous, no polling needed)
Low Quality Output
Problem: Image doesn't match expectations Solutions:
- Add more descriptive details to prompt
- Use
quality="high"in tool_config - Try iterative refinement
- Add style references ("photorealistic", "cinematic")
Dimensions Don't Match Sora
Problem: Generated image wrong size for video Solutions:
- Use
prepare_reference_imageafter download - Or specify size in tool_config:
tool_config=ImageGeneration( type="image_generation", size="1536x1024" # Matches Sora 1792x1024 better )
Content Filtered
Problem: "Moderation" error Solutions:
- Revise prompt to avoid policy violations
- Try
moderation="low"for borderline content - Check OpenAI usage policies
API Limits
- Rate limits: Varies by account tier
- Size limits: Input images must be reasonable size
- Timeout: Complex generations may take 30-60 seconds
- File formats: JPEG, PNG, WEBP supported
Next Steps
- Test basic generation: Start with simple prompts
- Experiment with refinement: Try iterative workflows
- Combine with video: Generate → Download → Prepare → Animate
- Explore advanced features: Image editing, masking, composition
See also:
- API Reference - Complete parameter documentation
- Reference Images - Using images with Sora
- Sora Prompting Guide - Video generation tips