Bird
Raised Fist0
Node.jsframework~8 mins

Handling worker crashes and restart in Node.js - Performance & Optimization

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Performance: Handling worker crashes and restart
HIGH IMPACT
This concept affects server responsiveness and stability, impacting how quickly the server recovers from worker failures without blocking requests.
Managing worker crashes in a Node.js cluster
Node.js
import cluster from 'node:cluster';
import os from 'node:os';

if (cluster.isPrimary) {
  const numCPUs = os.cpus().length;
  for (let i = 0; i < numCPUs; i++) {
    cluster.fork();
  }

  cluster.on('exit', (worker, code, signal) => {
    console.log(`Worker ${worker.process.pid} died, restarting...`);
    cluster.fork();
  });
} else {
  // Worker code
  import('./server.js');
}
Automatically restarts workers on crash, maintaining full capacity and responsiveness without manual intervention.
📈 Performance GainMaintains steady request handling capacity; reduces INP spikes caused by worker downtime.
Managing worker crashes in a Node.js cluster
Node.js
import cluster from 'node:cluster';
import os from 'node:os';

if (cluster.isPrimary) {
  const numCPUs = os.cpus().length;
  for (let i = 0; i < numCPUs; i++) {
    cluster.fork();
  }

  cluster.on('exit', (worker, code, signal) => {
    console.log(`Worker ${worker.process.pid} died`);
    // No restart logic here
  });
} else {
  // Worker code
  import('./server.js');
}
When a worker crashes, it is logged but not restarted, causing reduced capacity and slower response times until manual intervention.
📉 Performance CostBlocks handling of some requests until manual restart; increases INP due to fewer workers.
Performance Comparison
PatternWorker AvailabilityRequest DelayRecovery TimeVerdict
No restart on crashDecreases over timeIncreases significantlyManual restart needed[X] Bad
Automatic restart on crashMaintained at full capacityMinimal increaseImmediate recovery[OK] Good
Rendering Pipeline
In Node.js server clusters, worker crashes cause request handling delays. Restarting workers quickly restores capacity, minimizing request queuing and delays.
Request Handling
Event Loop
Process Management
⚠️ BottleneckReduced worker availability causes request queuing and slower event loop processing.
Core Web Vital Affected
INP
This concept affects server responsiveness and stability, impacting how quickly the server recovers from worker failures without blocking requests.
Optimization Tips
1Always listen to the 'exit' event to detect worker crashes.
2Restart workers immediately to maintain full concurrency and responsiveness.
3Avoid manual restarts to prevent prolonged request delays and degraded user experience.
Performance Quiz - 3 Questions
Test your performance knowledge
What is the main performance benefit of automatically restarting crashed workers in a Node.js cluster?
AMaintains server responsiveness by quickly restoring worker capacity
BReduces memory usage by killing workers permanently
CImproves CPU usage by limiting the number of workers
DPrevents any worker from ever crashing
DevTools: Node.js Inspector (Debugger) and Logs
How to check: Run the Node.js app with --inspect flag, monitor cluster worker exit events in console logs, and observe if workers restart automatically after crash.
What to look for: Look for logs showing worker crashes followed by immediate worker forks indicating automatic restarts, ensuring no prolonged request delays.

Practice

(1/5)
1. What is the main purpose of listening to the exit event on a worker in Node.js cluster module?
easy
A. To log the worker's CPU usage
B. To start a new worker automatically
C. To send messages between workers
D. To detect when a worker crashes or stops running

Solution

  1. Step 1: Understand the exit event role

    The exit event is triggered when a worker process stops, either normally or due to a crash.
  2. Step 2: Identify the purpose of listening to exit

    Listening to exit helps detect unexpected worker crashes so the master can respond.
  3. Final Answer:

    To detect when a worker crashes or stops running -> Option D
  4. Quick Check:

    exit event = detect crash [OK]
Hint: Remember: exit event means worker stopped or crashed [OK]
Common Mistakes:
  • Confusing exit event with message passing
  • Thinking exit event starts new workers automatically
  • Assuming exit event logs CPU usage
2. Which of the following is the correct way to listen for a worker's exit event in Node.js cluster?
easy
A. cluster.on('exit', worker => { /* handle exit */ });
B. worker.on('exit', () => { /* handle exit */ });
C. worker.listen('exit', () => { /* handle exit */ });
D. process.on('workerExit', () => { /* handle exit */ });

Solution

  1. Step 1: Recall event listener syntax on worker

    In Node.js cluster, each worker is an EventEmitter and uses on to listen to events.
  2. Step 2: Match correct event and method

    The correct event is exit and the method is on, so worker.on('exit', ...) is correct.
  3. Final Answer:

    worker.on('exit', () => { /* handle exit */ }); -> Option B
  4. Quick Check:

    Use on with exit on worker [OK]
Hint: Use worker.on('exit', callback) to catch exit events [OK]
Common Mistakes:
  • Using .listen instead of .on
  • Listening on cluster instead of worker
  • Using wrong event name like 'workerExit'
3. Given the code below, what will be logged when a worker crashes?
const cluster = require('cluster');
if (cluster.isMaster) {
  const worker = cluster.fork();
  worker.on('exit', (code, signal) => {
    console.log(`Worker exited with code ${code} and signal ${signal}`);
  });
} else {
  process.exit(1); // Simulate crash
}
medium
A. Worker exited with code null and signal SIGTERM
B. Worker exited with code 0 and signal null
C. Worker exited with code 1 and signal null
D. No output because exit event is not triggered

Solution

  1. Step 1: Understand process.exit(1) effect

    Calling process.exit(1) ends the worker with exit code 1, indicating an error.
  2. Step 2: Check exit event parameters

    The exit event callback receives the exit code and signal; here signal is null because no signal caused the exit.
  3. Final Answer:

    Worker exited with code 1 and signal null -> Option C
  4. Quick Check:

    Exit code 1 means crash, signal null if no signal [OK]
Hint: Exit code 1 means crash, signal null if no signal sent [OK]
Common Mistakes:
  • Assuming exit code 0 means crash
  • Confusing signal with exit code
  • Thinking exit event won't fire on crash
4. Identify the error in this code snippet that tries to restart a worker after it crashes:
const cluster = require('cluster');
if (cluster.isMaster) {
  cluster.fork();
  cluster.on('exit', (worker) => {
    console.log('Worker crashed, restarting...');
    cluster.fork();
  });
}
medium
A. The 'exit' event should be listened on 'cluster' but the callback parameter should be (worker, code, signal)
B. The 'exit' event should be listened on 'cluster.workers', not 'cluster'
C. The 'exit' event should be listened on 'cluster' but the callback parameters are wrong
D. The 'exit' event callback parameters are incorrect; it should receive (code, signal)

Solution

  1. Step 1: Check where to listen for worker exit

    The 'exit' event is emitted by the cluster module, and the callback receives (worker, code, signal).
  2. Step 2: Identify callback parameter mismatch

    The code uses only one parameter (worker), but the event provides three parameters; this can cause confusion or errors.
  3. Final Answer:

    The 'exit' event should be listened on 'cluster' but the callback parameter should be (worker, code, signal) -> Option A
  4. Quick Check:

    cluster.on('exit', (worker, code, signal)) is correct [OK]
Hint: cluster.on('exit') callback needs (worker, code, signal) parameters [OK]
Common Mistakes:
  • Listening on cluster.workers instead of cluster
  • Using wrong callback parameters
  • Ignoring code and signal parameters
5. You want to ensure your Node.js app automatically restarts a worker if it crashes, but only up to 3 restarts per minute to avoid infinite loops. Which approach best implements this behavior?
hard
A. Use a counter and timestamp in the master process to track restarts; restart only if under limit
B. Restart workers immediately on every exit event without limits
C. Use a setTimeout to delay restarts by 1 minute after each crash
D. Restart workers only if exit code is 0, ignore other exit codes

Solution

  1. Step 1: Understand the need to limit restarts

    Unlimited restarts can cause infinite loops if the worker crashes repeatedly.
  2. Step 2: Implement a counter and timestamp logic

    Track how many times workers restart within a time window (e.g., 3 restarts per minute) and only restart if under the limit.
  3. Final Answer:

    Use a counter and timestamp in the master process to track restarts; restart only if under limit -> Option A
  4. Quick Check:

    Limit restarts with counter and time check [OK]
Hint: Count restarts with time checks to avoid infinite loops [OK]
Common Mistakes:
  • Restarting without limits causing infinite loops
  • Delaying restarts but not counting attempts
  • Restarting only on exit code 0 (normal exit)