Chapter 18: Background Jobs & Task Queues: BullMQ, Redis Streams, Exponential Retries, and Worker Patterns
Chapter 18: Background Jobs & Task Queues: BullMQ, Cron Jobs, and Worker Patterns
HTTP is inherently synchronous: a client issues a request and waits on an open network socket for a response. In modern web applications, the accepted threshold for API response times is under 200 milliseconds.
However, many critical backend tasks are fundamentally slow or computationally intensive:
- Transcoding a user-uploaded 4K video into multiple HLS bitrates (minutes).
- Generating a 100-page accounting audit report in PDF format (5 to 15 seconds).
- Dispatching 5,000 transactional emails via a third-party API like SendGrid or Postmark (seconds to minutes).
- Syncing database state with external CRM webhooks.
If you attempt to execute these tasks directly inside an Express route handler, three disasters occur:
- HTTP Timeouts: Reverse proxies (Nginx, AWS ALB, Cloudflare) terminate the connection with a
504 Gateway Timeoutwhen responses take longer than 30–60 seconds. - Event Loop Starvation: CPU-intensive operations (PDF rendering, image compression, cryptographic hashing) block the single-threaded Node.js event loop, preventing all other concurrent requests from being processed.
- Catastrophic Data Loss: If you trigger asynchronous work as a background promise (
sendEmail().catch(...)) and the Node.js process crashes or restarts during a deployment, the in-memory task vanishes forever.
In this chapter, you will master the architecture of background job processing in Node.js:
- The Message Queue / Job Queue Pattern and asynchronous offloading.
- Why primitive in-memory approaches (
setTimeout, fire-and-forget Promises) fail in production. - BullMQ Architecture: Redis Streams, job states, and atomic Lua script coordination.
- Resilience Engineering: Exponential backoff retries, Dead Letter Queues (DLQ), and idempotent job design.
- Distributed Cron & Scheduled Jobs: Ensuring a cron runs exactly once across 20 container replicas.
- Separation of Concerns: Decoupling the HTTP API server from the Worker process with graceful shutdown (
SIGTERM). - Client Delivery Architecture: How the frontend retrieves completed artifacts (Polling vs. WebSockets/SSE vs. S3 Presigned URLs).
1. The Asynchronous Offloading Architecture
The solution to slow operations is to decouple the acceptance of work from its execution.
Real-World Analogy: The Restaurant Kitchen Ticket Wheel
Think of an asynchronous job queue like a busy restaurant:
- When you order a pizza at the front counter, the cashier doesn't run into the kitchen, knead dough, and bake the pizza for 15 minutes while a line of 40 customers waits behind you.
- Instead, the cashier prints a paper ticket (the Job Payload), clips it onto the revolving metal order carousel (the Persistent Queue), gives you an order number receipt (an HTTP 202 Accepted response), and instantly takes the next customer's order.
- In the kitchen, three line cooks (the Background Workers) continuously grab tickets from the carousel, cook the meals in parallel at a steady pace, and ring a bell when finished.
- Even if the cashier's computer reboots, the physical tickets on the kitchen carousel remain intact.
HTTP Request Lifecycle (Instant ~15ms):
Client ────────► POST /api/v1/reports/export (JSON)
◄──────── HTTP 202 Accepted { jobId: "job_9481", status: "queued" }
│
│ Producer pushes job payload
▼
┌───────────────────────────┐
│ Persistent Redis Queue │
│ (BullMQ Streams & Hashes)│
└─────────────┬─────────────┘
│
│ Consumer pulls job
▼
Background Worker Lifecycle (Decoupled, runs for 30s):
┌───────────────────────────┐
│ Dedicated Worker Process │
│ - Generates 100-page PDF │
│ - Uploads PDF to AWS S3 │
│ - Dispatches Email Link │
└───────────────────────────┘
Why Primitive "Fire-and-Forget" Anti-Patterns Fail
Many developers attempt to avoid queue infrastructure by simply omitting the await keyword:
// ❌ CRITICAL PRODUCTION ANTI-PATTERN: Fire-and-forget promise
app.post('/api/v1/orders', async (req, res) => {
const order = await orderService.createOrder(req.body);
// Firing async task in RAM without awaiting:
emailService.sendOrderReceipt(order.id).catch(console.error);
res.status(201).json({ success: true, order });
});
Why this breaks in production:
- Zero Crash Resilience: Node.js stores active Promises in V8 Heap memory. If Kubernetes terminates the pod, a rolling deployment occurs, or an uncaught exception restarts the container, the uncompleted promise is lost forever without any record.
- No Concurrency Limits: If 10,000 customers place an order simultaneously, your Node.js process attempts to spawn 10,000 concurrent network connections to your email provider, exhausting socket descriptors and triggering external API rate limit bans.
- No Retries with Backoff: If SendGrid experiences a 10-second network blip, the email fails permanently; there is no mechanism to retry 30 seconds later.
2. BullMQ: The Node.js Standard for Redis Queues
In the Node.js ecosystem, BullMQ is the leading library for background queues. BullMQ is written entirely in TypeScript and leverages modern Redis capabilities (Redis Streams, Hashes, and atomic Lua scripts) to guarantee that jobs are never duplicated, lost, or processed by multiple workers simultaneously.
pnpm add bullmq ioredis
The BullMQ Job Lifecycle
A job progresses through a deterministic state machine managed by Redis:
┌──────────────┐
│ Delayed │ (Scheduled or backoff retry)
└──────┬───────┘
│
▼
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Add ├─────────►│ Waiting ├─────────►│ Active │
│ (Producer) │ │ (In Queue) │ │ (In Worker) │
└──────────────┘ └──────────────┘ └──────┬───────┘
│
┌───────────────┴───────────────┐
│ │
▼ (Success) ▼ (Throws Error)
┌──────────────┐ ┌──────────────┐
│ Completed │ │ Failed │
└──────────────┘ └──────┬───────┘
│
┌──────────────┴──────────────┐
▼ (Has Retries Left) ▼ (Retries Exhausted)
┌──────────────┐ ┌──────────────┐
│ Backoff Wait │ │ Dead Letter │
│ (Delayed) │ │ Queue (DLQ) │
└──────────────┘ └──────────────┘
| State | Description |
|---|---|
| Waiting | Job is queued in Redis, waiting for an available worker thread. |
| Active | Job has been picked up by a worker and is currently executing. |
| Delayed | Job is scheduled to run at a specific future timestamp, or waiting on an exponential retry delay. |
| Completed | Job executed successfully without throwing an unhandled exception. |
| Failed | Job threw an error. If retries remain, it moves to Delayed; otherwise, it enters permanent failure. |
3. Production Architecture: Queues, Workers, and Types
Let's build a modular background processing architecture following our project conventions:
- Strongly typed job definitions using TypeScript interfaces.
- Decoupled
Queueproducer andWorkerconsumer. - Exponential backoff retry policies.
- Idempotency via deterministic
jobId.
1. Job Definitions & Types (src/features/jobs/jobs.types.ts)
// Define job names as an enum
export enum JobQueueName {
NOTIFICATION = 'notification-queue',
MEDIA_PROCESSING = 'media-processing-queue',
}
// Data payload for welcome emails
export interface WelcomeEmailJobData {
userId: string;
email: string;
name: string;
}
// Data payload for PDF report generation
export interface GenerateReportJobData {
reportId: string;
userId: string;
timeframe: 'monthly' | 'quarterly' | 'annual';
}
2. The Producer: Queue Manager (src/features/jobs/queue.service.ts)
The Queue instance is imported by your API controllers to add jobs:
import { Queue } from 'bullmq';
import Redis from 'ioredis';
import { JobQueueName, WelcomeEmailJobData, GenerateReportJobData } from './jobs.types';
// Shared Redis connection configuration for BullMQ
const redisConnection = new Redis(process.env.REDIS_URL || 'redis://localhost:6379', {
maxRetriesPerRequest: null, // Required by BullMQ
enableReadyCheck: false,
});
export class QueueService {
public static readonly notificationQueue = new Queue<WelcomeEmailJobData>(
JobQueueName.NOTIFICATION,
{
connection: redisConnection,
defaultJobOptions: {
// 🔒 Production Resilience:
attempts: 5, // Retry up to 5 times on transient errors
backoff: {
type: 'exponential',
delay: 3000, // Delays: 3s, 6s, 12s, 24s, 48s
},
removeOnComplete: { age: 86400, count: 5000 }, // Keep 5000 recent jobs for 24h
removeOnFail: { age: 604800 }, // Keep failed jobs for 7 days
},
}
);
public static readonly mediaQueue = new Queue<GenerateReportJobData>(
JobQueueName.MEDIA_PROCESSING,
{
connection: redisConnection,
defaultJobOptions: {
attempts: 3,
backoff: { type: 'exponential', delay: 5000 },
},
}
);
/**
* Adds an idempotent welcome email job.
*/
public static async dispatchWelcomeEmail(data: WelcomeEmailJobData) {
// 🔒 Idempotency Key: Prevents duplicate emails if API endpoint is retried
const jobId = `welcome-email:${data.userId}`;
return this.notificationQueue.add('send-welcome-email', data, {
jobId, // Unique ID ensures only ONE job exists for this key
});
}
/**
* Dispatches a background report generation job.
*/
public static async dispatchReportGeneration(data: GenerateReportJobData) {
return this.mediaQueue.add('generate-pdf-report', data);
}
}
4. The Consumer: Worker Process & Error Handling
Workers should run in a separate Node.js process (e.g. node dist/worker.js) from your web server (node dist/server.js). This ensures that heavy background computation never starves the HTTP event loop.
Implementing the Worker (src/worker.ts)
import { Worker, Job } from 'bullmq';
import Redis from 'ioredis';
import { JobQueueName, WelcomeEmailJobData } from './features/jobs/jobs.types';
import { logger } from './common/logger/logger';
const redisConnection = new Redis(process.env.REDIS_URL || 'redis://localhost:6379', {
maxRetriesPerRequest: null,
enableReadyCheck: false,
});
/**
* Worker handling transactional email notifications.
*/
const notificationWorker = new Worker<WelcomeEmailJobData>(
JobQueueName.NOTIFICATION,
async (job: Job<WelcomeEmailJobData>) => {
logger.info({ jobId: job.id, attempt: job.attemptsMade + 1 }, 'Processing notification job');
const { email, name } = job.data;
// Simulate external email API call (e.g., SendGrid/AWS SES)
await mockSendEmail(email, name);
logger.info({ jobId: job.id, email }, 'Welcome email sent successfully');
return { deliveredAt: new Date().toISOString() };
},
{
connection: redisConnection,
concurrency: 10, // Process up to 10 jobs concurrently per worker instance
limiter: {
max: 50, // Rate Limit: Max 50 emails
duration: 1000, // per 1,000 milliseconds (respecting external API quotas)
},
}
);
// Worker Lifecycle Event Handlers:
notificationWorker.on('completed', (job: Job) => {
logger.info({ jobId: job.id }, 'Job completed successfully');
});
notificationWorker.on('failed', (job: Job | undefined, err: Error) => {
if (job) {
logger.error(
{ jobId: job.id, attemptsMade: job.attemptsMade, err: err.message },
'Job failed execution'
);
// If all retries exhausted, send alert to Slack/PagerDuty!
if (job.attemptsMade >= (job.opts.attempts || 1)) {
logger.fatal({ jobId: job.id, data: job.data }, '🚨 JOB MOVED TO DEAD LETTER QUEUE (DLQ)');
}
}
});
// Helper simulating external email API
async function mockSendEmail(to: string, name: string): Promise<void> {
// Simulate 300ms network round-trip
await new Promise((resolve) => setTimeout(resolve, 300));
// Simulate occasional network timeout to demonstrate retry backoff
if (Math.random() < 0.05) {
throw new Error('503 Service Unavailable: Email Gateway Timeout');
}
}
// -------------------------------------------------------------
// GRACEFUL SHUTDOWN (Crucial for Kubernetes / Docker)
// -------------------------------------------------------------
const shutdown = async (signal: string) => {
logger.info(`Received ${signal}. Gracefully stopping worker...`);
// Closes worker: Waits for currently active jobs to finish without taking new ones
await notificationWorker.close();
await redisConnection.quit();
logger.info('Worker shutdown complete. Exiting.');
process.exit(0);
};
process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));
5. Distributed Cron & Scheduled Jobs
In traditional monolithic servers, developers frequently use node-cron or setInterval to trigger recurring jobs:
// ❌ DANGEROUS IN PRODUCTION CLUSTERS:
cron.schedule('0 0 * * *', () => {
generateDailyInvoices();
});
The Replica Collision Bug
If your application scales to 5 container replicas in Kubernetes, each container runs its own independent node-cron instance. At midnight, your invoice generation runs 5 times simultaneously, charging customers five times!
The BullMQ Solution: Distributed Repeatable Jobs
BullMQ stores recurring schedules in Redis using distributed locks. Regardless of whether you run 1 replica or 50 replicas, BullMQ guarantees that the cron executes exactly once across the entire cluster:
// src/features/jobs/cron.service.ts
import { QueueService } from './queue.service';
export class CronScheduleService {
/**
* Registers distributed recurring jobs in BullMQ.
*/
public static async registerRecurringJobs(): Promise<void> {
// Run nightly database cleanup every night at 2:00 AM UTC
await QueueService.mediaQueue.add(
'nightly-cleanup',
{ reportId: 'SYSTEM', userId: 'SYSTEM', timeframe: 'monthly' },
{
repeat: {
pattern: '0 2 * * *', // Standard 5-field cron syntax
},
jobId: 'system-nightly-cleanup', // Unique identifier
}
);
console.log('✅ Distributed cron job registered in Redis');
}
}
6. How the HTTP Controller Uses the Queue
Let's look at how clean our HTTP route handler becomes when offloading to BullMQ:
// src/features/user/user.controller.ts
import { Request, Response, NextFunction } from 'express';
import { QueueService } from '../jobs/queue.service';
import { ApiResponse } from '../../common/responses/api-response';
export class UserController {
public register = async (req: Request, res: Response, next: NextFunction): Promise<void> => {
try {
// 1. Fast, synchronous database transaction (saves user)
const newUser = await this.userService.createUser(req.body);
// 2. Asynchronous offloading (non-blocking, persistent in Redis)
await QueueService.dispatchWelcomeEmail({
userId: newUser.id,
email: newUser.email,
name: newUser.name,
});
// 3. Instant HTTP response (~20ms latency!)
ApiResponse.created(res, 'User registered successfully', newUser);
} catch (error) {
next(error);
}
};
}
7. The Result Dilemma: How Does the Frontend Get the Completed File?
A universal question engineers ask when adopting background queues is:
"I offloaded the PDF generation to BullMQ and sent an immediate
HTTP 202 Acceptedwith ajobId. The HTTP socket is now closed. How does the frontend actually get the generated PDF file? Does the client need to request the job again with the ID? What do production systems actually do?"
The Golden Rule: Never Store File Buffers in Redis
[!CAUTION] Never store raw binary buffers (PDFs, videos, ZIP archives) inside Redis or BullMQ job payloads! Redis is an in-memory database designed for ultra-low latency coordination. Storing 10MB PDF buffers in Redis will exhaust server RAM, degrade cluster performance, and trigger memory eviction policies.
Instead, production systems always follow the Offloaded Object Storage Pattern:
- The background worker renders the PDF (using Chromium/Puppeteer).
- The worker streams the file buffer directly to Object Storage (AWS S3, Cloudflare R2, or Google Cloud Storage).
- The worker generates a secure S3 Presigned URL (or stores the S3 file key in PostgreSQL).
- The worker returns lightweight metadata as its job completion result:JSON
{ "downloadUrl": "https://my-bucket.s3.amazonaws.com/reports/audit-2026.pdf?X-Amz-Signature=...", "fileName": "audit-2026.pdf", "fileSizeBytes": 482910 }
Now that the file lives safely in S3, here are the three production patterns used to deliver this download link to the user:
Pattern 1: Short Polling with Job ID (Industry Standard for 90% of Web Apps)
This is the exact pattern used by GitHub (when exporting repository archives), Stripe (when generating quarterly tax reports), and Linear.
1. Client POSTs request:
Client ────────────────────────► POST /api/v1/reports/pdf
◄──────────────────────── HTTP 202 Accepted { jobId: "job_8291", status: "queued" }
2. Client periodically polls status (every 2 seconds):
Client ────────────────────────► GET /api/v1/jobs/job_8291
◄──────────────────────── HTTP 200 OK { status: "active", progress: 45 }
Client ────────────────────────► GET /api/v1/jobs/job_8291
◄──────────────────────── HTTP 200 OK { status: "active", progress: 85 }
3. Job completes in S3; final poll returns download link:
Client ────────────────────────► GET /api/v1/jobs/job_8291
◄──────────────────────── HTTP 200 OK {
status: "completed",
result: { downloadUrl: "https://s3.amazonaws.com/..." }
}
Client automatically triggers browser download: window.location.href = downloadUrl
Step 1: Backend Job Status Controller (Express + BullMQ)
BullMQ provides built-in methods on the queue to inspect any job by its unique ID:
queue.getJob(jobId): Retrieves the job instance from Redis.await job.getState(): Returns'waiting' | 'active' | 'completed' | 'failed' | 'delayed'.job.progress: Reads current numerical progress reported by the worker viajob.updateProgress().job.returnvalue: Accesses the data returned by the worker upon completion.
// src/features/jobs/job-status.controller.ts
import { Request, Response, NextFunction } from 'express';
import { QueueService } from './queue.service';
import { ApiError } from '../../common/errors/api-error';
import { ApiResponse } from '../../common/responses/api-response';
export class JobStatusController {
/**
* GET /api/v1/jobs/:jobId
* Checks the real-time status of any background BullMQ task.
*/
public static async getStatus(req: Request, res: Response, next: NextFunction): Promise<void> {
try {
const { jobId } = req.params;
const job = await QueueService.mediaQueue.getJob(jobId);
if (!job) {
throw ApiError.notFound(`Job with ID ${jobId} was not found or has expired.`);
}
const state = await job.getState(); // 'waiting' | 'active' | 'completed' | 'failed'
if (state === 'completed') {
ApiResponse.ok(res, 'Job completed successfully', {
jobId: job.id,
status: 'completed',
result: job.returnvalue, // Contains S3 downloadUrl & metadata
finishedOn: job.finishedOn,
});
return;
}
if (state === 'failed') {
ApiResponse.ok(res, 'Job failed during execution', {
jobId: job.id,
status: 'failed',
failedReason: job.failedReason,
});
return;
}
// If still waiting or active, return current progress percentage
ApiResponse.ok(res, 'Job in progress', {
jobId: job.id,
status: state,
progress: job.progress || 0,
});
} catch (error) {
next(error);
}
}
}
Step 2: Frontend Implementation (React & TanStack React Query)
On the frontend, you don't need complicated setInterval hacks. Modern data-fetching libraries like TanStack React Query support smart polling via refetchInterval:
// src/features/reports/use-report-downloader.ts
import { useState, useEffect } from 'react';
import { useQuery } from '@tanstack/react-query';
export function useReportDownloader() {
const [jobId, setJobId] = useState<string | null>(null);
// Polls backend status endpoint every 2000ms until completed or failed
const { data: jobStatus, isFetching } = useQuery({
queryKey: ['job-status', jobId],
queryFn: async () => {
const res = await fetch(`/api/v1/jobs/${jobId}`);
if (!res.ok) throw new Error('Failed to fetch job status');
const json = await res.json();
return json.data;
},
enabled: Boolean(jobId), // Only start polling once we have a jobId
refetchInterval: (query) => {
const status = query.state.data?.status;
// Stop polling immediately once finished or failed!
if (status === 'completed' || status === 'failed') {
return false;
}
return 2000; // Poll every 2 seconds
},
});
// Automatically initiate file download when job finishes
useEffect(() => {
if (jobStatus?.status === 'completed' && jobStatus.result?.downloadUrl) {
const link = document.createElement('a');
link.href = jobStatus.result.downloadUrl;
link.setAttribute('download', jobStatus.result.fileName || 'report.pdf');
document.body.appendChild(link);
link.click();
link.remove();
}
}, [jobStatus]);
const requestPdf = async (reportId: string) => {
const res = await fetch('/api/v1/reports/pdf', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ reportId }),
});
const json = await res.json();
setJobId(json.data.jobId); // Triggers polling query
};
return { requestPdf, jobStatus, isPending: Boolean(jobId && jobStatus?.status !== 'completed') };
}
Pattern 2: Push Notifications via WebSockets or Server-Sent Events (SSE)
If the user is waiting on an interactive progress screen (e.g., rendering video clips or generating AI models) where polling every 2 seconds feels too slow or creates excessive HTTP requests, use Server-Sent Events (SSE) or WebSockets.
Worker completes job in Redis
│
│ redis.publish('job_events', { userId, jobId, downloadUrl })
▼
API Server (Redis Pub/Sub Subscriber)
│
│ Pushes event down open SSE stream or Socket.IO connection
▼
Client Browser (Receives real-time event instantly: zero polling!)
- When the client dispatches the request, it listens on an SSE stream (
GET /api/v1/events/stream) or joins a WebSocket room. - When the BullMQ worker completes the job, it publishes a completion message to a Redis Pub/Sub channel:TS
// Inside BullMQ Worker processor: const s3Url = await uploadPdfToS3(pdfBuffer); // Notify API servers via Redis Pub/Sub: await redisPublisher.publish( 'job_events', JSON.stringify({ userId: job.data.userId, jobId: job.id, type: 'PDF_READY', downloadUrl: s3Url, }) ); - Any running API instance listening to Redis receives the event and pushes it down the user's open SSE connection:TS
// In Frontend: const eventSource = new EventSource('/api/v1/events/stream'); eventSource.onmessage = (event) => { const data = JSON.parse(event.data); if (data.type === 'PDF_READY') { window.location.href = data.downloadUrl; } };
Pattern 3: Asynchronous Out-of-Band Delivery (Email / In-App Bell)
What if the PDF report is a 200-page enterprise financial audit that takes 3 to 5 minutes to compute?
Expecting a user to keep a browser tab open for 5 minutes is a terrible user experience. In production, applications like AWS Billing, Jira, and QuickBooks use Asynchronous Out-of-Band Delivery:
- Immediate User Feedback:
The API responds immediately:JSON
{ "status": "queued", "message": "Your annual audit report is being generated. Because this report takes several minutes, we will email you a secure download link as soon as it is ready." } - Worker Dispatches Email & Notification:
When the worker finishes generating and uploading the PDF to S3:
- It inserts a row into the database
notificationstable ({ userId, title: 'Your PDF is ready', link: s3PresignedUrl }). - It dispatches a transactional email (via Resend or SendGrid) containing a time-limited signed link:
"Your Annual Audit Report is ready. [Download Audit Report (Expires in 24 Hours)]"
- It inserts a row into the database
Production Comparison Matrix: Which Pattern Should You Choose?
| Architecture Pattern | Latency to User | Infrastructure Overhead | Best For |
|---|---|---|---|
Short Polling (/jobs/:id) |
1 – 2 seconds | Lowest (Standard REST, zero socket state, auto-retrying) | Standard tasks taking 2s to 20s (PDF generation, invoice downloads, CSV exports). |
| Server-Sent Events / WebSockets | Real-time (<50ms) | Medium (Requires maintaining persistent socket connections & Redis Pub/Sub) | Interactive UI with real-time progress bars (0% → 100%), video transcoding, AI image generation. |
| Email / In-App Notification Center | Minutes | Low (Decoupled from client session) | Heavy jobs taking >30 seconds (full database backups, complex enterprise audit reports). |
8. Production Background Jobs Checklist & Summary
┌────────────────────────────────────────────────────────────────────────────┐
│ PRODUCTION BACKGROUND JOBS CHECKLIST │
├────────────────────────────────────────────────────────────────────────────┤
│ [ ] No Fire-and-Forget Promises: All asynchronous work is backed by a │
│ persistent Redis queue; zero in-memory promise dangling. │
│ │
│ [ ] Process Isolation: Background workers execute in a dedicated Node.js │
│ process/container, completely isolated from HTTP web traffic. │
│ │
│ [ ] Exponential Backoff: Transient failures automatically retry with │
│ exponential delays to prevent hammering downstream third-party APIs. │
│ │
│ [ ] Dead Letter Queue (DLQ): Jobs that exhaust all retries trigger alerts │
│ and remain accessible for debugging and manual replay. │
│ │
│ [ ] Job Idempotency: Unique `jobId` keys prevent duplicate job creation │
│ when API clients retry requests. │
│ │
│ [ ] Rate Limiting on Workers: Worker concurrency and throughput configured │
│ to respect external third-party API rate quotas. │
│ │
│ [ ] Distributed Cron: Scheduled jobs managed via Redis repeat patterns │
│ to guarantee exactly-once execution across multi-container replicas. │
│ │
│ [ ] Graceful Worker Shutdown: `SIGTERM` handlers allow inflight jobs to │
│ finish or gracefully requeue before the container terminates. │
│ │
│ [ ] Artifact Storage Offloading: Raw binary files (PDFs, videos) are saved │
│ in Object Storage (S3/R2); Redis stores only metadata & signed URLs. │
│ │
│ [ ] Client Delivery Strategy: A clear polling, WebSocket, or notification │
│ flow defined so clients reliably receive completed artifacts. │
└────────────────────────────────────────────────────────────────────────────┘
In the next chapter, we will master Chapter 19: WebSockets & Real-Time Communication: Socket.IO, Namespaces, Redis Adapter, and SSE, establishing bi-directional communication channels and scaling them across multi-node clusters using the Redis Adapter.