Django deployment / production runbook

The Django Production Deployment Checklist: What manage.py check --deploy Misses

Go beyond static settings. Check database connections, reverse proxies, background jobs and migration risks before a release.

Interactive runbook · Local browser persistence · Updated September 28, 2026

Static checks vs. runtime reality

Why passing check --deploy does not guarantee uptime

Running python manage.py check --deploy is the first step recommended in the official Django docs. It inspects roughly 15 static settings: ensuring DEBUG = False, checking that your SECRET_KEY is not hardcoded, and checking secure cookie flags.

However, modern Django production runs in distributed environments: behind SSL-terminating reverse proxies (Nginx, AWS ALB, Cloudflare), connected to cloud databases with idle connection timeouts (AWS RDS, Supabase, Neon), packaged into Docker containers with asset manifests, and paired with asynchronous task queues (Celery). None of these operational interfaces can be validated by static code inspection.

Primary references: Django deployment checklist, Django database settings, and FetchNode production monitoring checklist.

Why check --deploy misses it

check --deploy does not evaluate database driver socket health or cloud timeout lifecycles.

# settings.py (Django 4.1+)
DATABASES = {
    "default": {
        "ENGINE": "django.db.backends.postgresql",
        # ...
        "CONN_MAX_AGE": 60,          # Reuse connection for up to 60 seconds
        "CONN_HEALTH_CHECKS": True,  # Test socket liveness before query execution
    }
}

Why check --deploy misses it

It cannot know your container replica count, Gunicorn worker/thread geometry, or PostgreSQL max_connections limit.

# settings.py (Django 5.1+ native pooling with psycopg 3)
DATABASES = {
    "default": {
        "ENGINE": "django.db.backends.postgresql",
        # ...
        "OPTIONS": {
            "pool": {
                "min_size": 1,
                "max_size": 4,
                "timeout": 10,
            }
        },
    }
}
# Note: If using PgBouncer in transaction mode, set CONN_MAX_AGE = 0.

To diagnose query volume and growth before pooling issues arise, explore our guide to fixing Django N+1 queries in production.

Why check --deploy misses it

check --deploy may verify CSRF_COOKIE_SECURE, but does not parse CSRF_TRUSTED_ORIGINS schemas or inspect whether the reverse proxy forwards the proto header.

# settings.py
# Only enable if you control the proxy and it strips spoofed client headers:
SECURE_PROXY_SSL_HEADER = ("HTTP_X_FORWARDED_PROTO", "https")

# Full URI schemes are mandatory in Django 4.0+:
CSRF_TRUSTED_ORIGINS = [
    "https://example.com",
    "https://www.example.com",
    "https://*.example.com",
]

Why check --deploy misses it

The linter recommends setting SECURE_SSL_REDIRECT = True blindly, without checking whether proxy communication is plain HTTP.

# nginx.conf
location / {
    proxy_pass http://django_upstream;
    proxy_set_header Host $host;
    proxy_set_header X-Real-IP $remote_addr;
    proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
    proxy_set_header X-Forwarded-Proto $scheme;  # Essential to prevent redirect loops
}

Why check --deploy misses it

Static manifests are generated only during collectstatic, not during setting checks.

# In Dockerfile: run collectstatic during image build, NOT at container startup
RUN python manage.py collectstatic --noinput

# In settings.py: prevent unhandled 500s if a third-party asset reference is missing
WHITENOISE_MANIFEST_STRICT = False

Why check --deploy misses it

check --deploy inspects Django core settings only and has no visibility into Celery or worker configurations.

# settings.py (with namespace="CELERY", sets task_time_limit)
CELERY_TASK_SOFT_TIME_LIMIT = 120    # Raises SoftTimeLimitExceeded inside task
CELERY_TASK_TIME_LIMIT = 150         # Sends SIGKILL to hung child worker process
CELERY_WORKER_PREFETCH_MULTIPLIER = 1  # Prevents long-running tasks from starving queues

# Always set explicit network timeouts in your tasks:
response = requests.get("https://api.example.com/webhook", timeout=(3.05, 30))

See our step-by-step diagnostic runbook on fixing stuck or hanging Celery tasks.

Why check --deploy misses it

It emits a generic warning (W018) suggesting you review logging, but provides no structured handler or scanner suppression.

# settings.py
LOGGING = {
    "version": 1,
    "disable_existing_loggers": False,
    "formatters": {
        "verbose": {"format": "%(levelname)s %(asctime)s %(module)s %(message)s"},
    },
    "handlers": {
        "console": {"class": "logging.StreamHandler", "formatter": "verbose"},
        "null": {"class": "logging.NullHandler"},
    },
    "root": {"handlers": ["console"], "level": "INFO"},
    "loggers": {
        "django.request": {"handlers": ["console"], "level": "ERROR", "propagate": False},
        "django.security.DisallowedHost": {"handlers": ["null"], "propagate": False},
    },
}

Learn more about capturing and grouping unhandled Django exceptions without scanner noise.

Why check --deploy misses it

Static checks cannot trace request execution paths or detect synchronous network I/O in view functions.

Use django-anymail with transactional HTTP APIs (Postmark, SES, Mailgun) or dispatch email via Celery background tasks so web workers return responses immediately.

# Option A: Transactional HTTP API via django-anymail (recommended)
# pip install django-anymail[postmark]
EMAIL_BACKEND = "anymail.backends.postmark.EmailBackend"
ANYMAIL = {
    "POSTMARK_SERVER_TOKEN": os.environ["POSTMARK_SERVER_TOKEN"],
}

# Option B: Asynchronous background task (Celery)
# send_welcome_email.delay(user_id=user.id)

Why check --deploy misses it

Migration safety depends on database table size, lock timeouts, and multi-phase deployment patterns.

# In PostgreSQL migrations:
from django.contrib.postgres.operations import AddIndexConcurrently
from django.db import migrations, models

class Migration(migrations.Migration):
    atomic = False  # Mandatory for concurrent index creation

    operations = [
        # 1. Fail fast if lock cannot be acquired within 2 seconds
        migrations.RunSQL("SET lock_timeout = '2s';"),
        # 2. Add index concurrently without locking the table
        AddIndexConcurrently(
            model_name="order",
            index=models.Index(fields=["created_at"], name="order_created_idx"),
        ),
    ]

Learn how query growth affects locking in our guide on diagnosing production database queries.

Why check --deploy misses it

Pre-deploy checks are a single point in time. Runtime issues evolve as traffic patterns shift, data grows, and external services degrade.

Deep dive

The four failure modes that cause silent production downtime

These four failure modes represent common causes of post-deployment emergency hotfixes across production Django environments.

Failure Mode 1 Cloud database connection drops

The 5-minute idle timeout crash

A developer deploys Django with CONN_MAX_AGE = 600 to save SSL negotiation overhead on PostgreSQL. Traffic is light overnight. A customer visits at 3:00 AM. In the intervening 15 minutes, AWS RDS or Supabase terminated the idle TCP socket silently. The Gunicorn worker issues a query over the dead socket. Instead of reconnecting, Django immediately crashes with OperationalError and displays a 500 page to the customer.

The permanent remedy:

Always pair CONN_MAX_AGE with CONN_HEALTH_CHECKS = True on Django 4.1+. If using a serverless driver or transaction pooler, keep CONN_MAX_AGE = 0.

Failure Mode 2 Reverse proxy CSRF mismatch

The 403 Forbidden on user login

Everything works locally, but immediately after deploying behind Nginx or AWS ALB, every login, registration, and checkout POST request returns 403 Forbidden: Origin checking failed. Developers frequently attempt setting CSRF_COOKIE_SECURE = False in frustration. The actual root cause: Django 4.0+ requires full scheme matching (https://example.com in CSRF_TRUSTED_ORIGINS) and requires SECURE_PROXY_SSL_HEADER = ('HTTP_X_FORWARDED_PROTO', 'https') so Django knows the incoming request was HTTPS.

Failure Mode 3 Docker Whitenoise missing manifest

The container startup 500 error

Using CompressedManifestStaticFilesStorage hashes file contents for aggressive browser caching. If a vendor CSS file contains a relative reference to a non-existent glyph (e.g. url('fonts/glyphicons.woff')), the manifest builder raises ValueError on page render. If collectstatic is executed at container entrypoint rather than image build time, containers crash in a restart loop in production.

Failure Mode 4 Synchronous email blocking workers

The worker freeze cascade

Using Django's built-in send_mail() inside a view makes a synchronous connection to your SMTP server. When a marketing email campaign drives 20 new signups in a minute, Gunicorn's 4 worker processes all stall waiting for SMTP handshakes. The server stops accepting any HTTP traffic, resulting in 504 Gateway Timeout errors across your entire portfolio.

Deployment runbook FAQ

Frequently asked deployment questions

What does manage.py check --deploy miss?

Django's manage.py check --deploy only validates static settings like DEBUG=False, SECRET_KEY, and cookie flags. It cannot verify runtime infrastructure: dropped idle connections on cloud databases, pool sizing, reverse proxy header forwarding for CSRF and SSL redirect loops, missing static asset manifests in Docker builds, unhandled worker timeouts in Celery, disappearing 500 error logs when SMTP is unconfigured, or table locks during database migrations.

Why does Django raise OperationalError: server closed the connection unexpectedly?

Cloud databases (AWS RDS, Supabase, Neon) and stateful NAT firewalls terminate idle TCP connections after 60 to 300 seconds. When CONN_MAX_AGE > 0 is set without CONN_HEALTH_CHECKS = True, Django reuses a dead TCP socket on the next request, resulting in an immediate crash. Setting CONN_HEALTH_CHECKS = True (Django 4.1+) instructs Django to test the connection before query execution, gracefully reconnecting if the socket died.

What is the difference between CONN_MAX_AGE and Django 5.1 connection pooling?

CONN_MAX_AGE keeps one persistent database connection per worker process or thread. With multi-process, multi-threaded Gunicorn setups, this can easily exhaust Postgres max_connections. Django 5.1 introduced native client-side connection pooling via psycopg 3 (DATABASES['OPTIONS']['pool']), allowing requests within a process to share a bounded connection pool. For large multi-container deployments, server-side poolers like PgBouncer or Supavisor remain standard.

Why does CSRF verification fail with 403 Forbidden behind reverse proxies?

In Django 4.0+, CSRF_TRUSTED_ORIGINS strictly requires explicit schemes (such as https://example.com, not just example.com). Additionally, when behind an SSL-terminating reverse proxy (Nginx, Traefik, AWS ALB, Cloudflare), you must set SECURE_PROXY_SSL_HEADER = ('HTTP_X_FORWARDED_PROTO', 'https') so Django knows the client initiated an HTTPS request rather than plain HTTP.

How do you prevent table locks during Django migrations?

Use multi-step expand-contract migrations: add new columns as nullable, deploy code that writes to both columns, backfill existing rows asynchronously, deploy code that reads from the new column, and finally remove the old column. For indexes, use PostgreSQL's AddIndexConcurrently with atomic = False, and configure lock_timeout = '2s' to prevent migration queries from blocking all live web traffic.

Move from static pre-deploy checks to continuous portfolio operations

FetchNode prioritizes which production Django website needs attention first across unhandled exceptions, slow queries, background tasks, and customer conversion funnels.

No credit card · Inspectable setup · Designed for Django operators maintaining 3–10 sites