Concurrency and Performance Guide
October 14, 2025 ยท View on GitHub
This document explains the concurrency limits and performance characteristics of MarkItDown Server.
๐ Current Concurrency Behavior
Default Configuration
By default, MarkItDown Server runs with:
- Single worker process (Uvicorn default)
- Async request handling via FastAPI/ASGI
- No explicit rate limiting
- Concurrent requests: Limited by system resources and async event loop
How It Works
The server uses FastAPI with Uvicorn, which provides:
- Asynchronous I/O: Multiple requests can be handled concurrently within a single worker
- Event loop: Non-blocking request handling
- File operations: Currently synchronous (blocking), but isolated per request
Performance Characteristics
With the default single-worker configuration:
- โ Good for: Low to moderate traffic (10-50 concurrent requests)
- โ Pros: Simple deployment, low memory footprint (~100-200MB)
- โ ๏ธ Limitations: CPU-bound operations (file conversion) can block other requests
- โ ๏ธ Not suitable for: High-traffic production without scaling
โ๏ธ Configuration Options
1. Multi-Worker Configuration
Run multiple worker processes to handle more concurrent requests:
# Set the number of workers via environment variable
export WORKERS=4
python app.py
Or with Docker:
docker run -d \
--name markitdownserver \
-p 8490:8490 \
-e WORKERS=4 \
ghcr.io/elbruno/markitdownserver:latest
Recommended worker count: (2 ร CPU_cores) + 1
2. Rate Limiting (Optional)
To prevent abuse, you can enable rate limiting:
# Limit to 10 requests per minute per IP
export ENABLE_RATE_LIMIT=true
export RATE_LIMIT="10/minute"
python app.py
3. Request Timeout
Configure timeout for long-running conversions:
# Set timeout to 300 seconds (5 minutes)
export REQUEST_TIMEOUT=300
๐ Production Deployment Recommendations
Small Scale (< 100 requests/minute)
Docker with resource limits:
docker run -d \
--name markitdownserver \
-p 8490:8490 \
-e WORKERS=2 \
--memory="512m" \
--cpus="1.0" \
ghcr.io/elbruno/markitdownserver:latest
Medium Scale (100-1000 requests/minute)
Kubernetes with auto-scaling:
apiVersion: apps/v1
kind: Deployment
metadata:
name: markitdownserver
spec:
replicas: 3
template:
spec:
containers:
- name: markitdownserver
image: ghcr.io/elbruno/markitdownserver:latest
env:
- name: WORKERS
value: "4"
resources:
requests:
memory: "512Mi"
cpu: "500m"
limits:
memory: "1Gi"
cpu: "1000m"
Large Scale (> 1000 requests/minute)
Azure Container Apps with auto-scaling:
az containerapp create \
--name markitdown-server \
--resource-group rg-markitdown \
--environment env-markitdown \
--image ghcr.io/elbruno/markitdownserver:latest \
--target-port 8490 \
--ingress external \
--min-replicas 2 \
--max-replicas 10 \
--cpu 1.0 \
--memory 2.0Gi \
--env-vars WORKERS=4
Auto-scaling triggers:
- Concurrent requests: 10 per replica
- CPU utilization: 70%
- Memory utilization: 80%
๐ Benchmarking
Testing Concurrency
Test your server's concurrency limits with Apache Bench:
# Install Apache Bench
sudo apt-get install apache2-utils # Ubuntu/Debian
brew install httpd # macOS
# Test with 100 concurrent requests (1000 total)
ab -n 1000 -c 100 http://localhost:8490/health
# Test file upload endpoint
ab -n 100 -c 10 -p test.pdf -T multipart/form-data \
http://localhost:8490/process_file
Load Testing with Locust
Create locustfile.py:
from locust import HttpUser, task, between
class MarkItDownUser(HttpUser):
wait_time = between(1, 3)
@task
def health_check(self):
self.client.get("/health")
@task(3)
def convert_file(self):
files = {"file": open("sample.pdf", "rb")}
self.client.post("/process_file", files=files)
Run load test:
pip install locust
locust -f locustfile.py --host=http://localhost:8490
๐ Security Considerations
Rate Limiting
Without rate limiting, the server is vulnerable to:
- DoS attacks: Overwhelming the server with requests
- Resource exhaustion: Large file uploads consuming memory
- Cost overruns: Excessive API usage in cloud deployments
Recommendation: Enable rate limiting in production:
export ENABLE_RATE_LIMIT=true
export RATE_LIMIT="60/minute"
File Size Limits
Current limit: 50MB per file
To adjust:
export MAX_FILE_SIZE=104857600 # 100MB in bytes
๐ Monitoring
Metrics to Monitor
- Request rate: Requests per second/minute
- Response time: p50, p95, p99 latency
- Error rate: 4xx and 5xx responses
- Resource usage: CPU, memory, disk I/O
- Concurrent requests: Active requests at any time
- Queue depth: Pending requests (if using queue)
Health Check
The /health endpoint provides basic health status:
curl http://localhost:8490/health
Response:
{
"status": "healthy",
"timestamp": "2025-10-13T23:45:18",
"service": "MarkItDown Server",
"version": "1.0.0"
}
Application Insights (Azure)
When deployed to Azure, integrate Application Insights:
export APPLICATIONINSIGHTS_CONNECTION_STRING="..."
โ FAQ
Q: How many concurrent requests can the server handle?
A: With default configuration (1 worker):
- Light load: 10-50 concurrent requests
- With 4 workers: 40-200 concurrent requests
- Actual limit: Depends on file size, conversion complexity, and system resources
Q: Does the server have a built-in concurrency limit?
A: No built-in hard limit by default. Limits are determined by:
- System resources (CPU, memory)
- Number of workers configured
- Optional rate limiting (if enabled)
- Deployment platform limits (e.g., Azure Container Apps)
Q: What happens when the limit is reached?
A: Requests will queue up and may experience:
- Increased response time
- Timeout errors (if conversion takes too long)
- HTTP 429 errors (if rate limiting is enabled)
- HTTP 503 errors (if server is overloaded)
Q: How can I increase the concurrency limit?
A: Multiple approaches:
- Increase worker count:
WORKERS=4 - Scale horizontally: Deploy multiple instances
- Optimize resources: Increase CPU/memory allocation
- Use async file operations (future enhancement)
Q: Is there a rate limit?
A: Not by default. You can enable optional rate limiting:
export ENABLE_RATE_LIMIT=true
export RATE_LIMIT="100/minute"
๐ฏ Best Practices
- Start with 2-4 workers for production
- Enable rate limiting to prevent abuse
- Monitor resource usage and adjust worker count
- Use load balancer for horizontal scaling
- Set appropriate timeouts for long conversions
- Configure health checks for auto-recovery
- Test under load before production deployment
๐ Future Enhancements
Planned improvements for better concurrency:
- Async file operations with aiofiles
- Redis-based request queue
- Response caching for identical files
- Streaming response for large files
- Worker pool for CPU-bound operations
- Advanced rate limiting with token bucket
- Circuit breaker for failing conversions
๐ Additional Resources
- FastAPI Deployment Guide
- Uvicorn Deployment
- Docker Production Best Practices
- Azure Container Apps Scaling
Need help? Open an issue on GitHub