Cloud & DevOpsMar 20268 min read

Celery + Redis for Async Processing: Patterns That Actually Scale

Celery is easy to adopt and easy to misuse. Idempotency, queue separation by SLA, late acknowledgment, and backpressure — the four patterns that separate production-grade task processing from a demo.

DILJOT SINGH · TECHNICAL LEAD

Every Django SaaS I've architected has ended up with Celery and Redis in it — report generation, webhook processing, email, data sync. Celery's strength is that a working setup takes an afternoon. Its danger is the same thing: the defaults are fine for a demo and quietly wrong for production. These are the patterns I now consider non-negotiable.

1. Every task is idempotent, no exceptions

Celery offers at-least-once delivery. Workers crash mid-task, visibility timeouts expire, retries fire — sooner or later, every task runs twice. If your task charges a card, sends an email, or increments a counter, duplicate execution is a user-visible bug.

The pattern: derive a deterministic idempotency key from the task's inputs, and check-and-set it before doing the side effect. For payment-adjacent work we stored keys in Postgres (same transaction as the business write); for lower-stakes work, Redis SETNX with a TTL is enough. The discipline matters more than the storage: a task that can't articulate its idempotency key isn't ready for production.

2. Queues separated by SLA, not by feature

The default single-queue setup means a burst of slow report-generation tasks delays the password-reset email behind them. The fix is queue separation — but the right axis is latency expectation, not feature area:

Three queues covered every system I've run. More than five and you're doing capacity planning by superstition.

3. Late acknowledgment, and honest time limits

task_acks_late = True
task_reject_on_worker_lost = True
task_time_limit = 600        # hard kill
task_soft_time_limit = 540   # raises SoftTimeLimitExceeded first
worker_prefetch_multiplier = 1  # for long-running queues

acks_late means a task is acknowledged after it completes, so a worker dying mid-task returns the task to the queue instead of losing it — this is why idempotency is rule one. Soft time limits give the task a chance to clean up before the hard kill. And prefetch_multiplier=1 on long-task queues stops one worker from hoarding ten tasks it won't reach for an hour.

4. Backpressure: watch queue depth, not worker CPU

The failure mode nobody alerts on: producers enqueue faster than workers drain, the Redis-backed queue grows silently, and by the time anyone notices, you're processing this morning's tasks at dinner time. Worker CPU looks healthy the whole time — they're busy, just hopelessly behind.

Redis deserves respect

Treat the broker as a stateful production dependency, not a cache you happen to enqueue into: separate the Celery broker from your application cache (different eviction policies — an LRU eviction eating queued tasks is a spectacular outage), turn on persistence if losing the queue on restart is unacceptable, and monitor Redis memory with the same seriousness as your database. Most "Celery is unreliable" stories I've debugged were Redis configuration stories wearing a costume.