HorizontalPodAutoscaler

Why the HPA is not scaling, and how to fix it

The HorizontalPodAutoscaler compares the pods’ metric with the target and adjusts the replica count. When it does not scale, it is almost always because it cannot read the metric, and it says so in its conditions and events.

1. See what the HPA sees

The TARGETS column shows the current metric and the target. <unknown> means it cannot read the metric:

kubectl get hpa -n <namespace>
kubectl describe hpa <name> -n <namespace>

What the conditions and events say

  • FailedGetResourceMetric or "unable to get metrics": the metrics-server is not installed or not responding. Check with kubectl top pods.
  • "missing request for cpu" (or memory): utilization is computed against the request, so every container in the pod needs a request for the measured resource.
  • ScalingLimited with TooManyReplicas: it reached maxReplicas. Raise the limit or review the target.
  • ScalingActive false: the HPA is disabled for that target, for example because the Deployment is at zero replicas.

2. Check the metrics

kubectl top pods -n <namespace>
kubectl get apiservice v1beta1.metrics.k8s.io

It scales, but not the way you expected

  • Slow to scale down: by default the HPA waits 5 minutes of low metrics before removing replicas (stabilization window), so it does not flap.
  • Ignores small changes: there is a 10% tolerance around the target.
  • It scales, but the pods stay Pending: there is no room on the nodes. The cluster autoscaler is missing or it hit its maximum node count.
  • A request set too high makes utilization look low and the HPA never scales; a request set too low makes it scale too early.

Without a terminal, in Kubepier

In Kubepier, the Horizontal Pod Autoscalers list shows each one’s metrics and targets, warning events such as FailedGetResourceMetric rise to the top and usage per namespace sits alongside, from any cluster, in the browser or on your phone.

Start for free