Bird
Raised Fist0
Node.jsframework~5 mins

Handling worker crashes and restart in Node.js

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Introduction

Sometimes worker processes stop working unexpectedly. Handling crashes and restarting workers helps keep your app running smoothly without downtime.

You run multiple worker processes to handle tasks in parallel.
You want your app to recover automatically if a worker crashes.
You need to keep your server stable and responsive.
You want to monitor worker health and restart them when needed.
Syntax
Node.js
import cluster from 'node:cluster';
import os from 'node:os';

if (cluster.isPrimary) {
  // Fork workers
  for (let i = 0; i < os.cpus().length; i++) {
    cluster.fork();
  }

  cluster.on('exit', (worker, code, signal) => {
    console.log(`Worker ${worker.process.pid} died. Restarting...`);
    cluster.fork();
  });
} else {
  // Worker code here
}

cluster.isPrimary checks if the current process is the main one that controls workers.

The exit event lets you detect when a worker stops and restart it.

Examples
Restart a worker immediately after it crashes.
Node.js
cluster.on('exit', (worker, code, signal) => {
  console.log(`Worker ${worker.process.pid} crashed.`);
  cluster.fork();
});
Restart a worker with a delay to avoid rapid crash loops.
Node.js
cluster.on('exit', (worker) => {
  setTimeout(() => {
    cluster.fork();
  }, 1000); // Restart after 1 second delay
});
Basic cluster setup with one worker and restart on crash.
Node.js
if (cluster.isPrimary) {
  cluster.fork();
  cluster.on('exit', (worker) => {
    console.log(`Worker ${worker.process.pid} died.`);
    cluster.fork();
  });
} else {
  // Worker code
}
Sample Program

This program starts one worker per CPU core. Each worker crashes after 2 seconds. The primary process detects the crash and restarts the worker automatically.

Node.js
import cluster from 'node:cluster';
import os from 'node:os';

if (cluster.isPrimary) {
  console.log(`Primary ${process.pid} is running`);

  // Fork workers equal to number of CPU cores
  for (let i = 0; i < os.cpus().length; i++) {
    cluster.fork();
  }

  cluster.on('exit', (worker, code, signal) => {
    console.log(`Worker ${worker.process.pid} died. Restarting...`);
    cluster.fork();
  });
} else {
  console.log(`Worker ${process.pid} started`);

  // Simulate a crash after 2 seconds
  setTimeout(() => {
    console.log(`Worker ${process.pid} crashing now.`);
    process.exit(1);
  }, 2000);
}
OutputSuccess
Important Notes

Always monitor worker crashes to avoid infinite restart loops.

You can add logging or alerts inside the exit event handler for better monitoring.

Use process.exit(code) in workers to simulate crashes during testing.

Summary

Use the cluster module to run multiple workers for better performance.

Listen to the exit event to detect worker crashes.

Restart workers automatically to keep your app running smoothly.

Practice

(1/5)
1. What is the main purpose of listening to the exit event on a worker in Node.js cluster module?
easy
A. To log the worker's CPU usage
B. To start a new worker automatically
C. To send messages between workers
D. To detect when a worker crashes or stops running

Solution

  1. Step 1: Understand the exit event role

    The exit event is triggered when a worker process stops, either normally or due to a crash.
  2. Step 2: Identify the purpose of listening to exit

    Listening to exit helps detect unexpected worker crashes so the master can respond.
  3. Final Answer:

    To detect when a worker crashes or stops running -> Option D
  4. Quick Check:

    exit event = detect crash [OK]
Hint: Remember: exit event means worker stopped or crashed [OK]
Common Mistakes:
  • Confusing exit event with message passing
  • Thinking exit event starts new workers automatically
  • Assuming exit event logs CPU usage
2. Which of the following is the correct way to listen for a worker's exit event in Node.js cluster?
easy
A. cluster.on('exit', worker => { /* handle exit */ });
B. worker.on('exit', () => { /* handle exit */ });
C. worker.listen('exit', () => { /* handle exit */ });
D. process.on('workerExit', () => { /* handle exit */ });

Solution

  1. Step 1: Recall event listener syntax on worker

    In Node.js cluster, each worker is an EventEmitter and uses on to listen to events.
  2. Step 2: Match correct event and method

    The correct event is exit and the method is on, so worker.on('exit', ...) is correct.
  3. Final Answer:

    worker.on('exit', () => { /* handle exit */ }); -> Option B
  4. Quick Check:

    Use on with exit on worker [OK]
Hint: Use worker.on('exit', callback) to catch exit events [OK]
Common Mistakes:
  • Using .listen instead of .on
  • Listening on cluster instead of worker
  • Using wrong event name like 'workerExit'
3. Given the code below, what will be logged when a worker crashes?
const cluster = require('cluster');
if (cluster.isMaster) {
  const worker = cluster.fork();
  worker.on('exit', (code, signal) => {
    console.log(`Worker exited with code ${code} and signal ${signal}`);
  });
} else {
  process.exit(1); // Simulate crash
}
medium
A. Worker exited with code null and signal SIGTERM
B. Worker exited with code 0 and signal null
C. Worker exited with code 1 and signal null
D. No output because exit event is not triggered

Solution

  1. Step 1: Understand process.exit(1) effect

    Calling process.exit(1) ends the worker with exit code 1, indicating an error.
  2. Step 2: Check exit event parameters

    The exit event callback receives the exit code and signal; here signal is null because no signal caused the exit.
  3. Final Answer:

    Worker exited with code 1 and signal null -> Option C
  4. Quick Check:

    Exit code 1 means crash, signal null if no signal [OK]
Hint: Exit code 1 means crash, signal null if no signal sent [OK]
Common Mistakes:
  • Assuming exit code 0 means crash
  • Confusing signal with exit code
  • Thinking exit event won't fire on crash
4. Identify the error in this code snippet that tries to restart a worker after it crashes:
const cluster = require('cluster');
if (cluster.isMaster) {
  cluster.fork();
  cluster.on('exit', (worker) => {
    console.log('Worker crashed, restarting...');
    cluster.fork();
  });
}
medium
A. The 'exit' event should be listened on 'cluster' but the callback parameter should be (worker, code, signal)
B. The 'exit' event should be listened on 'cluster.workers', not 'cluster'
C. The 'exit' event should be listened on 'cluster' but the callback parameters are wrong
D. The 'exit' event callback parameters are incorrect; it should receive (code, signal)

Solution

  1. Step 1: Check where to listen for worker exit

    The 'exit' event is emitted by the cluster module, and the callback receives (worker, code, signal).
  2. Step 2: Identify callback parameter mismatch

    The code uses only one parameter (worker), but the event provides three parameters; this can cause confusion or errors.
  3. Final Answer:

    The 'exit' event should be listened on 'cluster' but the callback parameter should be (worker, code, signal) -> Option A
  4. Quick Check:

    cluster.on('exit', (worker, code, signal)) is correct [OK]
Hint: cluster.on('exit') callback needs (worker, code, signal) parameters [OK]
Common Mistakes:
  • Listening on cluster.workers instead of cluster
  • Using wrong callback parameters
  • Ignoring code and signal parameters
5. You want to ensure your Node.js app automatically restarts a worker if it crashes, but only up to 3 restarts per minute to avoid infinite loops. Which approach best implements this behavior?
hard
A. Use a counter and timestamp in the master process to track restarts; restart only if under limit
B. Restart workers immediately on every exit event without limits
C. Use a setTimeout to delay restarts by 1 minute after each crash
D. Restart workers only if exit code is 0, ignore other exit codes

Solution

  1. Step 1: Understand the need to limit restarts

    Unlimited restarts can cause infinite loops if the worker crashes repeatedly.
  2. Step 2: Implement a counter and timestamp logic

    Track how many times workers restart within a time window (e.g., 3 restarts per minute) and only restart if under the limit.
  3. Final Answer:

    Use a counter and timestamp in the master process to track restarts; restart only if under limit -> Option A
  4. Quick Check:

    Limit restarts with counter and time check [OK]
Hint: Count restarts with time checks to avoid infinite loops [OK]
Common Mistakes:
  • Restarting without limits causing infinite loops
  • Delaying restarts but not counting attempts
  • Restarting only on exit code 0 (normal exit)