Explorer
Node.js

Chapter 17: Caching & Performance: Redis, Cache-Aside Pattern, Stampede Prevention, and HTTP ETags

Chapter 17: Caching & Performance: Redis, Cache-Aside Pattern, Stampede Prevention, and HTTP ETags

In any scalable backend architecture, the database is the primary bottleneck.

A relational database like PostgreSQL or a document store like MongoDB must manage disk I/O, maintain index trees (B-Trees), enforce transactional isolation (ACID), and coordinate row-level locks. Under heavy concurrent load, complex JOIN queries or multi-stage aggregations will quickly saturate CPU cores and connection pools, causing response times to balloon from 20ms to 2,000ms.

Caching is the art and science of saving expensive computational results in fast, volatile memory so subsequent identical requests can be served in sub-millisecond time.

In this chapter, you will master the end-to-end caching architecture of production Node.js backends:

  1. The Multi-Tier Caching Hierarchy: L1 (Process Memory), L2 (Distributed Redis), and L3 (HTTP/Edge/CDN).
  2. Core Caching Patterns: Cache-Aside (Lazy Loading), Write-Through, and Write-Behind.
  3. The Three Production Caching Disasters: Cache Stampede (Dog-Piling), Cache Penetration, and Cache Avalanche—and how to eliminate them.
  4. HTTP Caching: Cache-Control, ETags, and conditional 304 Not Modified workflows.
  5. Production Redis Architecture: Serialization, key namespacing, and building a type-safe CacheService with stampede-proof mutex locks.
  6. Advanced Redis Data Structures: Beyond key-value strings to Hashes and Sorted Sets.

1. The Multi-Tier Caching Hierarchy

Not all caches are created equal. High-throughput distributed backends utilize a tiered caching strategy, placing the fastest caches closest to the execution runtime.

Real-World Analogy: The Office Information Flow

Think of this hierarchy like working in an engineering office:

  • L1 (In-Memory Heap) is the sticky note stuck directly on your computer monitor. It takes 1 microsecond to glance at, but it only holds a few notes and only you can see it.
  • L2 (Redis Cluster) is the large whiteboard in the central hallway. Any engineer on the floor can walk over, read it, or update it in a couple of seconds.
  • Tier 0 (Primary Database) is the underground file archive in the basement. It stores every document permanently with strict safety locks, but fetching a file requires filling out a requisition form, riding the elevator down, and searching through steel filing cabinets.
TEXT
Incoming HTTP Request
        │
        ▼
┌─────────────────────────────────────────────────────────────┐
│ Tier 3: HTTP / Edge Cache (CDN / Cloudflare / Browser)      │
│  - Distance: Edge Network                                   │
│  - Latency: ~10ms - 30ms (Never touches your backend!)      │
│  - Mechanism: Cache-Control, ETag, 304 Not Modified         │
└──────────────────────────────┬──────────────────────────────┘
                               │ (Cache Miss)
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ Tier 1: In-Memory L1 Cache (Node.js Heap LRU / Map)         │
│  - Distance: Inside the Node.js V8 Process Memory           │
│  - Latency: ~0.001ms (Nanoseconds!)                         │
│  - Limitation: Not shared across cluster replicas           │
└──────────────────────────────┬──────────────────────────────┘
                               │ (Cache Miss)
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ Tier 2: Distributed L2 Cache (Redis Cluster)                │
│  - Distance: Dedicated In-Memory Cache Cluster              │
│  - Latency: ~0.5ms - 2ms (TCP Round Trip)                   │
│  - Advantage: Shared state across 50+ Node.js replicas      │
└──────────────────────────────┬──────────────────────────────┘
                               │ (Cache Miss)
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ Tier 0: Primary Database (PostgreSQL / MongoDB)             │
│  - Distance: Persistent Disk & Buffer Pool                  │
│  - Latency: 15ms - 250ms+                                   │
└─────────────────────────────────────────────────────────────┘

Comparing Caching Tiers

Tier Technology Latency Scope Best Used For
L1 (In-Memory) Node.js Memory (lru-cache) ~1 µs Single Process Extremely hot configuration, static metadata, compiled templates.
L2 (Distributed) Redis / Dragonfly ~1 ms Entire Application Cluster User sessions, problem details, leaderboard rankings, API responses.
L3 (HTTP / Edge) Fastly / Cloudflare / Browser ~10–30 ms Global Edge Public static assets, read-only catalog feeds, public course pages.

2. Core Caching Patterns

The lifecycle of how data moves into and out of your cache defines your application's consistency and reliability.

1. The Cache-Aside Pattern (Lazy Loading)

Cache-Aside is the most widely adopted caching strategy for REST APIs. The application code sits between the cache and the database:

  1. Application receives a request for resource X.
  2. Application checks the cache:
    • Cache HIT: Return cached data immediately.
    • Cache MISS: Query database, store result in cache with a TTL (Time-To-Live), and return data.
TEXT
                     Cache-Aside Workflow

                      1. GET /problems/two-sum
                      ────────────────────────►
              Client                            Node.js API
                      ◄────────────────────────
                        4. Return Data
                              │          ▲
            2. Check Cache    │          │ 3b. Save in Cache (TTL: 1h)
                              ▼          │
                        ┌──────────────────┐
                        │   Redis Cache    │
                        └──────────────────┘
                              ▲
                              │ 3a. Read DB on Miss
                              ▼
                        ┌──────────────────┐
                        │    PostgreSQL    │
                        └──────────────────┘
  • Pros: Only requested data is cached (lazy loading); resilient to cache failure (if Redis goes down, requests still succeed by falling back to the database).
  • Cons: Cache miss penalty (the initial request pays the latency of both Redis lookup and DB query); requires explicit cache invalidation when data changes.

2. Write-Through & Write-Behind

  • Write-Through: The application writes to the cache and the database simultaneously in a single transaction. The cache is always up to date, eliminating cold misses, but write latency is higher.
  • Write-Behind (Write-Back): The application writes immediately to the cache and returns a 200 OK to the client. An asynchronous background worker batches the writes to PostgreSQL every few seconds.
    • Risk: If the Redis instance crashes before writes are persisted to PostgreSQL, data is permanently lost. Used primarily for high-throughput counters (e.g. video view counts, page impression analytics).

3. The Three Production Caching Disasters

Implementing caching naively without addressing concurrency edge cases will inevitably trigger production outages.

Disaster 1: The Cache Stampede (Dog-Piling / Thundering Herd)

A Cache Stampede occurs when a heavily requested piece of data (e.g., the homepage problem list receiving 2,000 requests/sec) expires from the cache.

TEXT
Time 00:00:00 ──► Cache Key `homepage:problems` expires!
Time 00:00:01 ──► 2,000 concurrent requests arrive at the same millisecond.
             ──► ALL 2,000 requests encounter a CACHE MISS simultaneously!
             ──► ALL 2,000 requests execute identical heavy SQL queries on PostgreSQL!
             ──► PostgreSQL Connection Pool Exhausted (100% CPU) 💥 DATABASE CRASH!

Real-World Analogy: The Coffee Pot Stampede

Imagine an office kitchen with a single coffee pot. At 9:00 AM, the pot runs dry. If 50 coworkers walk in simultaneously and all see an empty pot, imagine all 50 people trying to plug in 50 grinders and brew fresh beans in the exact same coffee machine at once. The kitchen breaker trips and the kitchen shuts down.

Instead, the first person to arrive grabs the "Barista Badge" (the distributed mutex lock), announces "I'm brewing a fresh pot!", and everyone else waits 60 seconds at their desk until the fresh pot is ready.

The Solution: Distributed Mutex Locking (Single-Flight)

When a cache miss occurs, the first request acquires an atomic distributed lock in Redis with a random token using SET lock:key token PX 5000 NX. Only that single request executes the heavy database query and repopulates the cache. All other requests wait briefly or receive the stale value.

Critically, the lock must be released using a Lua script to verify ownership, preventing a slow request from deleting another process's newly acquired lock:

TS
import { randomUUID } from 'node:crypto';

const lockToken = randomUUID();
// 1. Atomic lock acquisition with unique ownership token:
const acquired = await redis.set(`lock:${cacheKey}`, lockToken, 'PX', 5000, 'NX');

if (acquired) {
  try {
    const freshData = await queryPostgres();
    await redis.set(cacheKey, JSON.stringify(freshData), 'EX', 3600);
  } finally {
    // 2. Safe atomic release via Lua script (verifies token ownership)
    const releaseScript = `
      if redis.call("get", KEYS[1]) == ARGV[1] then
        return redis.call("del", KEYS[1])
      else
        return 0
      end
    `;
    await redis.eval(releaseScript, 1, `lock:${cacheKey}`, lockToken);
  }
} else {
  // Wait 50ms and retry reading from cache
  await new Promise((res) => setTimeout(res, 50));
  return await redis.get(cacheKey);
}

Disaster 2: Cache Penetration

Cache Penetration occurs when an attacker requests keys that do not exist in the database (e.g., GET /api/v1/problems/non-existent-id-99999).

  • Because the record does not exist in the database, the application never caches anything.
  • Every subsequent request for that non-existent ID bypasses the cache completely and hammers PostgreSQL with queries.

The Solution: Cache Null Objects with Short TTL

When the database returns null, store null or a sentinel value in Redis with a short expiration (e.g., 60 seconds):

TS
if (!dbResult) {
  await redis.set(cacheKey, JSON.stringify(null), 'EX', 60);
  return null;
}

Disaster 3: Cache Avalanche

A Cache Avalanche happens when hundreds of different cache keys are configured with the exact same TTL (e.g., EX: 3600). If they were all seeded at deployment time, they all expire at the exact same second, causing thousands of database queries simultaneously.

The Solution: Randomized TTL Jitter

Never use static expiration times. Always add randomized jitter:

TS
// Add random jitter of ±10% to prevent synchronized expiration
const baseTtl = 3600; // 1 hour
const jitter = Math.floor(Math.random() * 300); // 0-300 seconds
const finalTtl = baseTtl + jitter;

4. HTTP Caching: ETags and Conditional 304 Requests

The fastest request is the one your server never has to transmit.

HTTP specifies the Entity Tag (ETag) mechanism (RFC 9110). An ETag is an opaque identifier (often an MD5 hash of the response body or a version timestamp) assigned to a resource representation.

TEXT
First Request:
  Client  ────────►  GET /api/v1/problems/two-sum
  Client  ◄────────  200 OK
                     ETag: "w/38a9-f81d4"
                     Content-Length: 40960 (40KB sent)

Subsequent Request (Validating freshness):
  Client  ────────►  GET /api/v1/problems/two-sum
                     If-None-Match: "w/38a9-f81d4"
  Client  ◄────────  304 Not Modified
                     Content-Length: 0 (ZERO body bytes sent!)

Why 304 Not Modified is a Superpower

  1. 99% Bandwidth Savings: When a resource hasn't changed, the server sends only headers and an empty body.
  2. Instant Client Rendering: The browser reuses the cached response from disk instantly.

Implementing Fast ETag Checks with Version Timestamps

Instead of computing an MD5 hash across a massive JSON payload on every request, check the record's updatedAt timestamp directly:

TS
// src/features/problem/problem.controller.ts
export async function getProblemBySlug(req: Request, res: Response) {
  const { slug } = req.params;

  // 1. Fetch lightweight metadata first (or from cache)
  const meta = await db.query.problems.findFirst({
    where: eq(problems.slug, slug),
    columns: { id: true, updatedAt: true },
  });

  if (!meta) {
    throw ApiError.notFound('Problem not found');
  }

  // 2. Generate fast ETag from updatedAt timestamp
  const eTag = `"${meta.id}-${new Date(meta.updatedAt).getTime()}"`;
  res.setHeader('ETag', eTag);
  res.setHeader('Cache-Control', 'public, max-age=60, stale-while-revalidate=300');

  // 3. Conditional validation: If client ETag matches, return 304 immediately!
  if (req.headers['if-none-match'] === eTag) {
    res.status(304).end();
    return;
  }

  // 4. Only fetch full details (description, test cases) on cache miss
  const fullProblem = await problemService.getProblemDetails(meta.id);
  ApiResponse.ok(res, 'Problem fetched', fullProblem);
}

5. Production Redis Architecture: The CacheService

Let's build a production-grade CacheService using ioredis. It features:

  • Strongly typed get/set operations with automatic JSON serialization.
  • Dynamic TTL with randomized jitter.
  • Stampede-proof getOrSet() method using distributed Redis mutex locking.
  • Pattern-based cache invalidation.
TS
// src/common/cache/cache.service.ts
import Redis from 'ioredis';
import { logger } from '../logger/logger';

export interface CacheOptions {
  ttlSeconds?: number;
  jitterSeconds?: number;
}

export class CacheService {
  private readonly client: Redis;

  constructor() {
    this.client = new Redis(process.env.REDIS_URL || 'redis://localhost:6379', {
      maxRetriesPerRequest: 3,
      enableReadyCheck: true,
      retryStrategy: (times) => Math.min(times * 100, 3000),
    });

    this.client.on('error', (err) => {
      logger.error({ err }, 'Redis connection error');
    });

    this.client.on('connect', () => {
      logger.info('Connected to Redis Cache cluster');
    });
  }

  /**
   * Retrieves an item from cache, parsing JSON safely.
   */
  public async get<T>(key: string): Promise<T | null> {
    try {
      const data = await this.client.get(key);
      if (!data) return null;
      return JSON.parse(data) as T;
    } catch (error) {
      logger.warn({ error, key }, 'Cache read failed. Falling back to DB.');
      return null;
    }
  }

  /**
   * Sets an item in cache with TTL and random jitter to avoid cache avalanches.
   */
  public async set<T>(key: string, value: T, options: CacheOptions = {}): Promise<void> {
    try {
      const baseTtl = options.ttlSeconds ?? 3600; // Default 1 hour
      const maxJitter = options.jitterSeconds ?? 180; // 0 to 180s jitter
      const jitter = Math.floor(Math.random() * maxJitter);
      const finalTtl = baseTtl + jitter;

      const serialized = JSON.stringify(value);
      await this.client.set(key, serialized, 'EX', finalTtl);
    } catch (error) {
      logger.error({ error, key }, 'Failed to write to cache.');
    }
  }

  /**
   * Deletes a specific key or invalidates all keys matching a glob pattern.
   */
  public async del(key: string): Promise<void> {
    try {
      await this.client.del(key);
    } catch (error) {
      logger.error({ error, key }, 'Failed to delete cache key.');
    }
  }

  /**
   * Stampede-Proof Cache-Aside:
   * Retrieves key from cache. On miss, uses an atomic distributed lock
   * to ensure only ONE worker hits the database, preventing dog-piling.
   */
  public async getOrSet<T>(
    key: string,
    fetcher: () => Promise<T>,
    options: CacheOptions = {}
  ): Promise<T> {
    // 1. Check cache first
    const cached = await this.get<T>(key);
    if (cached !== null) {
      return cached;
    }

    // 2. Cache miss: Attempt to acquire mutex lock (5s lease time) with unique token
    const lockKey = `lock:${key}`;
    const lockToken = Math.random().toString(36).slice(2) + Date.now().toString(36);
    const acquiredLock = await this.client.set(lockKey, lockToken, 'PX', 5000, 'NX');

    if (!acquiredLock) {
      // Another concurrent request is already fetching from DB!
      // Wait 100ms and check cache again:
      await new Promise((resolve) => setTimeout(resolve, 100));
      const retryCached = await this.get<T>(key);
      if (retryCached !== null) {
        return retryCached;
      }
    }

    try {
      // 3. This process acquired the lock: execute DB fetcher
      const freshData = await fetcher();

      // 4. Save result in cache
      await this.set(key, freshData, options);

      return freshData;
    } finally {
      // 5. Release distributed lock safely via Lua script
      if (acquiredLock) {
        const releaseLua = `
          if redis.call("get", KEYS[1]) == ARGV[1] then
            return redis.call("del", KEYS[1])
          else
            return 0
          end
        `;
        await this.client.eval(releaseLua, 1, lockKey, lockToken);
      }
    }
  }
}

Using CacheService in a Problem Repository

TS
// src/features/problem/problem.service.ts
export class ProblemService {
  constructor(
    private readonly db: DatabaseClient,
    private readonly cache: CacheService
  ) {}

  public async getProblemById(id: string) {
    const cacheKey = `problem:${id}`;

    // Stampede-proof cache-aside lookup:
    return this.cache.getOrSet(
      cacheKey,
      async () => {
        const record = await this.db.query.problems.findFirst({
          where: eq(problems.id, id),
        });

        if (!record) {
          throw ApiError.notFound('Problem not found');
        }
        return record;
      },
      { ttlSeconds: 1800 } // 30 minutes + randomized jitter
    );
  }

  public async updateProblem(id: string, updateData: any) {
    const updated = await this.db.update(problems).set(updateData).where(eq(problems.id, id));

    // 🔒 Invalidate cache immediately upon update!
    await this.cache.del(`problem:${id}`);

    return updated;
  }
}

6. Advanced Redis Data Structures for Performance

Redis is not just a key-value string store. Using specialized Redis data structures avoids expensive database aggregations and in-memory sorting:

1. Hashes (HSET, HGETALL) for Objects

Instead of serializing an entire user session object as a JSON string, store it as a Redis Hash. This allows reading or updating a single field without deserializing the whole object:

TS
// Update only the lastActive timestamp without touching profile data
await redis.hset(`session:${sessionId}`, 'lastActive', Date.now());

2. Sorted Sets (ZADD, ZREVRANGE) for Leaderboards

Calculating user rankings across 100,000 users in PostgreSQL requires a slow ORDER BY points DESC LIMIT 10.

In Redis, a Sorted Set keeps elements ordered by score in memory with O(log(N)) complexity:

TS
// Increment points when user solves a problem:
await redis.zincrby('leaderboard:global', 50, userId);

// Fetch Top 10 users in microseconds:
const topUsers = await redis.zrevrange('leaderboard:global', 0, 9, 'WITHSCORES');

3. Redis Memory Limits & Eviction Policies

Redis is an in-memory data store. If your dataset grows larger than available RAM, the host operating system's Out-Of-Memory (OOM) killer will terminate the Redis process.

To prevent this, production Redis instances are configured with explicit memory limits and eviction algorithms in redis.conf:

CONF
# Limit Redis memory footprint to 4 Gigabytes
maxmemory 4gb

# Evict the least recently used keys when maxmemory is reached
maxmemory-policy allkeys-lru

Official Redis Eviction Policies Comparison

Policy Behavior on Memory Saturation Recommended Use Case
noeviction (Default) Rejects write commands with an Out-of-Memory error (OOM command not allowed). Read operations continue to work. When Redis is used as a primary database or message broker where data loss is unacceptable.
allkeys-lru Evicts the Least Recently Used keys across the entire dataset, regardless of whether a TTL was set. The standard choice for general-purpose API caching. Popular items stay hot; cold items fade away.
volatile-lru Evicts the Least Recently Used keys, but only among keys that have an expiration (TTL) defined. When the same Redis instance mixes persistent data (no TTL) and cache data (with TTL).
allkeys-lfu Evicts the Least Frequently Used keys. Tracks access frequency rather than just recency. When temporal bursts should not immediately displace consistently popular long-term keys.
volatile-ttl Evicts keys with the shortest remaining Time-To-Live. When expiring short-lived items first aligns best with business logic.

7. Production Caching Checklist & Summary

CODE
┌────────────────────────────────────────────────────────────────────────────┐
│                    PRODUCTION BACKEND CACHING CHECKLIST                    │
├────────────────────────────────────────────────────────────────────────────┤
│ [ ] Multi-Tiered Architecture: L1 in-memory for micro-lookups, L2 Redis    │
│     cluster for shared state, L3 HTTP/CDN for static and public media.     │
│                                                                            │
│ [ ] Stampede Defense: Hot queries protected with distributed mutex locks   │
│     or probabilistic early expiration (`getOrSet`).                        │
│                                                                            │
│ [ ] Randomized TTL Jitter: All cache expirations add randomized offsets    │
│     to prevent synchronized cache avalanches.                              │
│                                                                            │
│ [ ] Penetration Defense: Null database results cached with short TTL to    │
│     prevent non-existent ID queries from reaching the database.            │
│                                                                            │
│ [ ] Event-Driven Invalidation: Mutations (`PUT`, `DELETE`) invalidate or   │
│     refresh corresponding cache keys immediately.                          │
│                                                                            │
│ [ ] HTTP ETag Optimization: Dynamic resources send lightweight ETags;     │
│     conditional `304 Not Modified` eliminates unneeded bandwidth.          │
│                                                                            │
│ [ ] Connection Resilience: Redis client configured with bounded retries    │
│     and graceful fallback to DB when the cache is temporarily unavailable. │
└────────────────────────────────────────────────────────────────────────────┘

In the next chapter, we will master Chapter 18: Background Jobs & Task Queues: BullMQ, Cron Jobs, and Worker Patterns, learning how to offload heavy asynchronous work (sending emails, video transcoding, PDF generation, webhooks) so our HTTP responses remain blazing fast.

Finished this lesson?

Mark this chapter complete to update your learning streak and unlock the next lesson.