Failing pod

How to fix OOMKilled in Kubernetes

OOMKilled means the container used more memory than its limit and the kernel killed the process. The pod restarts, sometimes ends up in CrashLoopBackOff, and the application log shows no error, because the app had no time to write one.

1. Confirm the OOMKilled

Under Last State, describe shows Reason: OOMKilled and Exit Code: 137:

kubectl describe pod <pod> -n <namespace>
kubectl get pod <pod> -n <namespace> \
  -o jsonpath='{.status.containerStatuses[*].lastState.terminated.reason}'

2. Compare usage with the limit

kubectl top needs the metrics-server in the cluster:

kubectl top pod <pod> -n <namespace> --containers
kubectl get pod <pod> -n <namespace> \
  -o jsonpath='{.spec.containers[*].resources}'

Container limit or node memory

  • If the container went over its own limits.memory, the OOMKilled is its own: the limit is too low for what the app uses, or the app is leaking memory.
  • If the node ran out of memory, the kubelet evicts pods (Evicted status, memory pressure event), starting with those furthest over their request. Requests set too low let the scheduler pack too many pods on one node.

How to fix it

  • Set limits.memory above the real peak, with headroom, and requests.memory close to normal usage.
  • Runtimes with their own heap must respect the limit: on the JVM, use -XX:MaxRAMPercentage; on Node.js, --max-old-space-size; the .NET GC already reads the container limit.
  • If usage keeps growing until it hits the limit, raising it only delays the crash: it is a memory leak in the app.

Without a terminal, in Kubepier

In Kubepier, CPU and memory usage shows per cluster, node and namespace, from the metrics-server the cluster already has, and the pod restarted by OOMKilled rises to the failing pods list, in the browser or on your phone.

Start for free