⚡ ~/naveed/lab DevOps Lab
Home > DevOps Lab > Tutorials & Runbooks > Linux SRE Incident Runbook
Incident Response Linux Kernel SRE On-Call Runbook

Linux SRE Production Incident Runbook: Triaging OOMKilled, High I/O Wait & Packet Drops

1. The 60-Second Initial Triage Script

When PagerDuty wakes you up at 3:00 AM, avoid jumping into random logs. Execute Brendan Gregg's classic Linux USE method commands in order:

bash — 60-second health check
# 1. System load & CPU core count comparison
uptime

# 2. Kernel panic, OOM killer, or hardware events
dmesg -T | grep -Ei "oom|kill|segfault|error|reset" | tail -n 25

# 3. CPU vs I/O blocked threads (check r and b columns)
vmstat 1 5

# 4. Available RAM vs Buffers/Cache vs Swap
free -m

# 5. Disk utilization and await latency
iostat -xz 1 3

# 6. Socket counts and connection queue state
ss -s

2. Scenario 1: Triaging OOMKilled Processes (Out of Memory)

Symptom: Microservice or container suddenly terminates with exit code 137. Kubernetes pods show Reason: OOMKilled.

Run these commands on the affected Linux host or worker node:

bash — identify culprit process
# View the exact process killed and its memory score
dmesg -T | grep -A 10 -B 2 -i "invoked oom-killer"

# Top 15 memory-consuming processes right now
ps -eo pid,ppid,cmd,%mem,%cpu --sort=-%mem | head -n 15

# Inspect cgroup memory usage (cgroups v2)
cat /sys/fs/cgroup/memory.current
cat /sys/fs/cgroup/memory.events

Immediate Mitigation:

  • If on Kubernetes: Increase pod resources.limits.memory or debug JVM/Go heap allocation flags.
  • If on bare Linux host: Tune sysctl vm.overcommit_memory=0 or set a negative oom_score_adj for mission-critical daemons (e.g., echo -500 > /proc/<PID>/oom_score_adj).

3. Scenario 2: High Load Average with Low CPU (I/O Wait & D-State)

High load average does not necessarily mean high CPU usage! A Linux load average of 45 on an 8-core CPU with only 12% CPU usage indicates processes are stuck in Uninterruptible Sleep (D state) waiting for disk I/O, NFS mounts, or locked kernel syscalls.

bash — find processes in D state
# Catch all D-state processes
ps -eo stat,pid,user,cmd | grep "^D"

# Track which PID is doing massive disk read/write throughput
pidstat -d 1 5

# Inspect stack trace of a stuck process
cat /proc/<STUCK_PID>/wchan
cat /proc/<STUCK_PID>/stack

4. Scenario 3: Packet Drops & TCP Socket Exhaustion

Under heavy traffic, packets can be dropped at the NIC level, the kernel socket backlog, or the TCP accept queue.

bash — check socket and queue drops
# Check listen queue overflows (backlog full)
netstat -s | grep -i "listen"

# Look at Recv-Q vs Send-Q on listening sockets
# When Recv-Q == Send-Q on a LISTEN socket, the application is failing to accept() fast enough
ss -lnt

# Check NIC driver ring buffer drops
ethtool -S eth0 | grep -Ei "drop|discard|err"