Production Deployment Checklist
July 14, 2026 · View on GitHub
Time: ~15 minutes | Level: DevOps
Use this checklist before deploying ASAP agents to production. Covers security, monitoring, and scaling.
Prerequisites: Building Your First Agent, Building Resilient Agents
Security
TLS and HTTPS
- Use HTTPS for all production endpoints (no HTTP)
- Configure valid SSL certificates (not self-signed)
- Enable TLS 1.2+ (1.3 recommended)
- Set HSTS header:
Strict-Transport-Security: max-age=31536000; includeSubDomains - Redirect HTTP to HTTPS at reverse proxy (nginx, traefik, etc.)
Authentication
- Enable authentication in manifest (
auth=AuthScheme(schemes=["bearer"])or["basic"]) - Implement and wire
token_validatortocreate_app() - Store tokens in environment variables (never hardcode)
- Rotate tokens regularly; use short-lived tokens (15–60 min)
- Ensure clients send
Authorizationheader (Bearer or Basic) - If mounting
/usage,/sla, or/auditbeyond localhost:require_operator_auth=True+oauth2_config
Rate Limiting
- Configure rate limits via
ASAP_RATE_LIMITorrate_limitincreate_app() - Tune limits for your workload (e.g.
10/second;100/minute) - For multi-instance deployments, set
ASAP_RATE_LIMIT_BACKEND=redis://...(requirespip install 'asap-protocol[redis]'; transport uses thelimitslibrary) - Monitor
asap_rate_limit_exceeded_total(or equivalent) for abuse
Request Size Limits
- Set
max_request_sizeorASAP_MAX_REQUEST_SIZE(default 10MB) - Align with reverse proxy (
client_max_body_size) and ASGI server limits - Monitor 413 (Payload Too Large) responses
Handler Security
- Validate all payloads with Pydantic models (e.g.
TaskRequest) - Do not send unknown keys on
TaskRequest.config/CommonMetadata(forbidden since v2.5.2) - Use
FilePart(or equivalent) for file URIs; block path traversal (../) - Use
sanitize_for_logging()before logging envelope/payload - Never log secrets or PII
- Use
validate_handler()or equivalent before registering handlers
Replay Protection
- Keep timestamp validation enabled (default)
- Consider
require_nonce=Truefor high-security flows - Use a shared nonce store (Redis, DB) for multi-instance deployments
- Multi-worker Host/Agent JWT: inject
RedisJtiReplayCacheviaidentity_jti_cache(and MCPjti_replay_cache)
Monitoring
Metrics
- Expose metrics at
/asap/metrics(Prometheus format) - Configure Prometheus to scrape
/asap/metrics - Track:
asap_requests_total,asap_requests_error_total,asap_request_duration_seconds - Add custom metrics for business logic (e.g. tasks per skill)
Logging
- Use structured logging (
configure_logging,get_logger) - Set
ASAP_LOG_FORMAT=jsonfor production - Bind
trace_idandcorrelation_idfor distributed tracing - Ensure logs are aggregated (e.g. ELK, Loki, CloudWatch)
Health and Readiness
- Implement
/health(liveness) and/ready(readiness) endpoints - Wire readiness to dependencies (DB, Redis, downstream agents)
- Configure Kubernetes liveness/readiness probes
- Test graceful shutdown (SIGTERM) and drain in-flight requests
Alerting
- Alert on high error rate (e.g. >10% for 5 min)
- Alert on high latency (e.g. P99 > 1s)
- Alert on agent down (scrape failures)
- Alert on rate limit hits above threshold
Scaling
Client Configuration
- Use connection pooling (ASAPClient default; tune
pool_connections,pool_maxsizeif needed) - Enable compression for payloads > 1KB (default in ASAPClient)
- Configure
RetryConfig(retries, circuit breaker) for resilience - Set appropriate timeouts (
timeout,MANIFEST_REQUEST_TIMEOUT)
Server Configuration
- Use multiple uvicorn workers or Gunicorn with Uvicorn workers
- Configure worker count based on CPU (e.g. 2–4 per core)
- Set
DEFAULT_POOL_MAXSIZEandDEFAULT_POOL_CONNECTIONSfor high concurrency
Caching
- Enable manifest caching (
ManifestCache) for downstream agent manifests - Tune cache TTL (e.g. 5 min) for your discovery pattern
- Use shared cache (Redis) for multi-instance deployments if needed
State and Snapshots
- Use a persistent
SnapshotStore(Redis, PostgreSQL) for long-running tasks - Avoid
InMemorySnapshotStorein production (data lost on restart) - Tune checkpoint frequency; keep snapshots lean
Horizontal Scaling
- Run agents as stateless pods/containers where possible
- Use sticky sessions or shared state (Redis) if required
- Configure autoscaling (HPA, cluster autoscaler) based on CPU/memory/request rate
Resilience
- Enable retries with backoff on the client
- Enable circuit breaker for failing downstream agents
- Implement fallbacks for critical paths (cached value, backup agent)
- Use state snapshots for resumable long-running tasks
Quick Reference
| Area | Key docs |
|---|---|
| Security | Security Guide |
| Metrics | Metrics Guide |
| Transport | Transport Guide |
| Resilience | Building Resilient Agents |