Incident Response
Linux Kernel SRE
On-Call Runbook
Linux SRE Production Incident Runbook: Triaging OOMKilled, High I/O Wait & Packet Drops
1. The 60-Second Initial Triage Script
When PagerDuty wakes you up at 3:00 AM, avoid jumping into random logs. Execute Brendan Gregg's classic Linux USE method commands in order:
bash — 60-second health check
# 1. System load & CPU core count comparison
uptime
# 2. Kernel panic, OOM killer, or hardware events
dmesg -T | grep -Ei "oom|kill|segfault|error|reset" | tail -n 25
# 3. CPU vs I/O blocked threads (check r and b columns)
vmstat 1 5
# 4. Available RAM vs Buffers/Cache vs Swap
free -m
# 5. Disk utilization and await latency
iostat -xz 1 3
# 6. Socket counts and connection queue state
ss -s
2. Scenario 1: Triaging OOMKilled Processes (Out of Memory)
Symptom: Microservice or container suddenly terminates with exit code
137. Kubernetes pods show Reason: OOMKilled.
Run these commands on the affected Linux host or worker node:
bash — identify culprit process
# View the exact process killed and its memory score
dmesg -T | grep -A 10 -B 2 -i "invoked oom-killer"
# Top 15 memory-consuming processes right now
ps -eo pid,ppid,cmd,%mem,%cpu --sort=-%mem | head -n 15
# Inspect cgroup memory usage (cgroups v2)
cat /sys/fs/cgroup/memory.current
cat /sys/fs/cgroup/memory.events
Immediate Mitigation:
- If on Kubernetes: Increase pod
resources.limits.memoryor debug JVM/Go heap allocation flags. - If on bare Linux host: Tune
sysctl vm.overcommit_memory=0or set a negativeoom_score_adjfor mission-critical daemons (e.g.,echo -500 > /proc/<PID>/oom_score_adj).
3. Scenario 2: High Load Average with Low CPU (I/O Wait & D-State)
High load average does not necessarily mean high CPU usage! A Linux load average of 45 on an 8-core CPU with only 12% CPU usage indicates processes are stuck in Uninterruptible Sleep (D state) waiting for disk I/O, NFS mounts, or locked kernel syscalls.
bash — find processes in D state
# Catch all D-state processes
ps -eo stat,pid,user,cmd | grep "^D"
# Track which PID is doing massive disk read/write throughput
pidstat -d 1 5
# Inspect stack trace of a stuck process
cat /proc/<STUCK_PID>/wchan
cat /proc/<STUCK_PID>/stack
4. Scenario 3: Packet Drops & TCP Socket Exhaustion
Under heavy traffic, packets can be dropped at the NIC level, the kernel socket backlog, or the TCP accept queue.
bash — check socket and queue drops
# Check listen queue overflows (backlog full)
netstat -s | grep -i "listen"
# Look at Recv-Q vs Send-Q on listening sockets
# When Recv-Q == Send-Q on a LISTEN socket, the application is failing to accept() fast enough
ss -lnt
# Check NIC driver ring buffer drops
ethtool -S eth0 | grep -Ei "drop|discard|err"
⚡ Explore Ecosystem Resources