Kubernetes HPA
Kubernetes HPA
Beginner
Q1: What is Kubernetes HPA?
HPA (Horizontal Pod Autoscaler) automatically scales the number of Pods in a workload based on observed resource usage or custom metrics.
Q2: Why do we need autoscaling?
Autoscaling helps applications handle variable load without manual intervention.
Q3: What is horizontal scaling?
Horizontal scaling adds more Pod replicas to handle more traffic or work.
Q4: What is vertical scaling?
Vertical scaling changes the resource size of a Pod, such as increasing CPU or memory.
Q5: Why is HPA called horizontal autoscaling?
Because it scales the number of replicas horizontally.
Q6: What kinds of workloads does HPA work with?
HPA typically works with Deployments, ReplicaSets, StatefulSets, and similar scalable workloads.
Q7: What is a target utilization?
A target utilization is the desired percentage of resource usage that HPA tries to maintain.
Q8: What is CPU utilization in HPA?
CPU utilization is the ratio of current CPU usage to the requested CPU for a Pod.
Q9: What is a target CPU utilization?
It is the CPU percentage HPA aims to achieve for the workload.
Q10: What is average utilization?
Average utilization is the average of resource usage across Pods in a scalable workload.
Q11: What is a minReplicaCount?
The minimum number of Pods allowed by the HPA.
Q12: What is a maxReplicaCount?
The maximum number of Pods allowed by the HPA.
Q13: What is a currentReplicaCount?
The current number of Pods in the workload.
Q14: What is desiredReplicaCount?
The target number of Pods that HPA intends to scale the workload to.
Q15: What is the HPA controller?
The HPA controller continuously evaluates metrics and updates the replica count.
Q16: Why is HPA implemented as a controller?
Because it continuously reconciles desired scale with observed metrics.
Q17: What is scale-out?
Scale-out means increasing Pod replicas to handle more load.
Q18: What is scale-in?
Scale-in means decreasing Pod replicas when demand drops.
Q19: Why is scaling important for cost and performance?
It helps keep services responsive while avoiding excessive wasted capacity.
Q20: What is a Deployment in Kubernetes?
A Deployment manages a ReplicaSet and maintains a desired number of replicas.
Q21: Why does HPA usually target a Deployment?
Because a Deployment is the common workload pattern for stateless apps.
Q22: What does HPA update?
HPA updates the replica count of the target workload.
Q23: How does HPA decide to scale?
It compares observed metrics to targets and calculates the desired replica count.
Q24: What is metric aggregation?
Metric aggregation combines values from multiple Pods to compute a workload-level metric.
Q25: What is a metrics server?
The Metrics Server collects resource metrics like CPU and memory for pods and nodes.
Q26: What is a custom metric?
A custom metric is a business or application-specific metric, such as queue depth or request rate.
Q27: What is an external metric?
An external metric originates outside the cluster, such as a cloud service or queue in another system.
Q28: Why do custom and external metrics matter?
Because CPU-only scaling is not always enough for real workload decisions.
Q29: What is an HPA target type?
The target type describes what kind of metric or resource HPA uses, such as CPU utilization or a custom metric.
Q30: What is a `deployment` or `HorizontalPodAutoscaler` resource?
It is the Kubernetes object representing the HPA configuration.
Q31: What is `apiVersion: autoscaling/v2`?
It is a newer HPA API version used for richer scaling behavior and multiple metrics.
Q32: Why is autoscaling useful with microservices?
Because different services experience different traffic patterns and need different scale behavior.
Q33: Why is having a stable pod request important?
HPA relies on CPU request and other metrics to estimate whether the workload is under or over capacity.
Q34: Why are resource requests critical to HPA?
Because HPA compares current utilization to target utilization based on workload requests.
Q35: What is a Pod request for CPU?
It is the amount of CPU the pod expects to use and that the scheduler reserves for it.
Q36: Why is an adequate CPU request important for HPA?
If the request is too low, HPA may think the pod is “overutilized” or scale unpredictably.
Q37: Why is a too-high CPU request a problem?
It can make HPA scale too conservatively or create unschedulable capacity issues.
Q38: Why do poor requests cause poor scaling?
Because HPA computes utilization based on request values and observed usage.
Q39: What is a scaling event?
A scaling event is a change to the desired pod count triggered by metrics.
Q40: Why do scaling events matter?
They directly affect cost, availability, and response time.
Q41: What is average CPU usage?
The average CPU used across the replicas in the workload.
Q42: What is CPU utilization formula?
Utilization is roughly current CPU / requested CPU * 100%.
Q43: Why does HPA compare to target utilization?
It seeks to maintain the workload near the target value to balance responsiveness and cost.
Q44: What is the default HPA behavior?
The default HPA behavior is to scale based on CPU utilization when configured.
Q45: What is a target value for custom metrics?
The target value is the desired metric value HPA tries to maintain.
Q46: What is a metric average value?
For some metric types, HPA calculates the average across Pods or a selected set of workloads.
Q47: What is `behavior` in HPA?
Behavior defines scaling policies for scaling up and down such as stabilization windows and policy rules.
Q48: Why is scaling behavior important?
It prevents rapid flapping or oscillation as metrics move.
Q49: What is `scaleUp` policy?
It defines how aggressive HPA should be when scaling up.
Q50: What is `scaleDown` policy?
It defines how conservative HPA should be when scaling down.
Q51: What is a stabilization window?
A stabilization window prevents HPA from reacting too quickly to short-term metric spikes.
Q52: Why is stabilization important?
It reduces churn and makes scale decisions more stable.
Q53: What is HPA `selectPolicy`?
It chooses how HPA picks the desired replicas under multiple scaling metrics.
Q54: What is `maxReplicas` and `minReplicas`?
The minimum and maximum replicas allowed in the HPA range.
Q55: Why do HPA policies matter in production?
Because aggressive scaling can cause thrash or poor cost control.
Q56: What is demand-based scaling?
It increases or decreases replicas based on actual traffic or workload.
Q57: What is a scale-down delay?
It is a delay or stabilization period that prevents quick scale-in after a brief spike.
Q58: What is a scale-up delay?
It is a delay or stabilization policy used to avoid too-frequent scale-out events.
Q59: Why do scale-up and scale-down differ?
Because scaling out often needs to be faster than scaling in, especially for latency-sensitive applications.
Q60: What is a custom metric example?
Examples include queue length, requests per second, or a business metric like active users.
Q61: What is an external metric example?
A message queue depth in a cloud service or a third-party API throughput measure.
Q62: What is the role of metrics adapters?
Metrics adapters provide custom and external metrics to the HPA controller.
Q63: Why are adapters needed?
Because HPA does not directly know every custom metric source.
Q64: What is Prometheus Adapter?
Prometheus Adapter is a common adapter that exposes Prometheus metrics as custom metrics for HPA.
Q65: Why is Prometheus popular for HPA?
Because it is widely used for application-level metrics and system observability.
Q66: What is a queue depth metric?
Queue depth is the number of pending tasks in a queue, which can indicate demand.
Q67: Why use queue depth for HPA?
Because CPU may not reflect actual demand for async processing workloads.
Q68: What is a request-rate metric?
It measures incoming requests per second and can be a strong indicator for web workloads.
Q69: Why is queue depth better than CPU for some workloads?
Because some workloads are I/O or queue-based and CPU usage may not track demand well.
Q70: Why is application-level scaling sometimes better than CPU scaling?
Because CPU utilization may not correlate to user-visible demand or business capacity.
Q71: What is a scaling threshold?
A threshold is a metric value that triggers scale-out or scale-in.
Q72: What is a pod count formula?
HPA calculates a new desired replica count using the target metric and current metric.
Q73: Why should pods have requests set before using HPA?
Without requests, CPU utilization calculations may not make sense or can be misleading.
Q74: What is a HPA cooldown?
A cooldown is a delay period before another scaling action occurs.
Q75: What does `apiVersion: autoscaling/v2beta2` mean?
It is an older HPA API version; newer clusters typically use v2 or v2beta2 depending on version.
Q76: What is a target value for CPU percent?
Often 50, 60, or 70 percent is used.
Q77: Why is 80% a common target?
Because it balances utilization and cushion but varies by workload type.
Q78: What is the risk of using too-high target utilization?
It may allow pods to saturate or become slow before scaling happens.
Q79: What is the risk of using too-low target utilization?
It may over-scale and increase costs without enough gain.
Q80: Why is scaling policy tuning important?
It avoids oscillation or poor cost/performance balance.
Q81: What is a `behaviour` policy example?
For example: scaleUp in steps, scaleDown slower than scaleUp.
Q82: Why do scaleDown policies often differ from scaleUp?
Because scale-down is usually more conservative to avoid thrashing or disruption.
Q83: What is `selectPolicy: Max` in HPA?
It chooses the maximum of multiple metric calculations when multiple metrics are used.
Q84: What is `selectPolicy: Min`?
It chooses the minimum required replica count under multiple metrics.
Q85: Why do multiple metrics matter?
Because some workloads may be limited by CPU, memory, or queue depth at different times.
Q86: What is a custom metric threshold used for HPA?
It defines the target queue depth, latency, or throughput value.
Q87: What is a multi-metric HPA?
A multi-metric HPA bases decisions on more than one metric.
Q88: Why is multi-metric HPA useful?
Because real workloads may need several signals, such as CPU and queue depth.
Q89: What is a workload-specific metric?
A metric unique to the application or service, like business transactions per second.
Q90: Why is `kubectl get hpa` important?
It lets you check current HPA status and scaling decisions.
Q91: What is HPA status in Kubernetes?
It includes current metrics, desired replicas, and the scaling event history.
Q92: Why is HPA status valuable?
It gives insight into whether autoscaling is functioning as expected.
Q93: What are common reasons HPA does not scale?
- missing metrics-server
- no resource requests
- custom metrics not configured
- low target or wrong thresholds
- pods not ready
- workload not scalable
Q94: Why is missing metrics-server a common issue?
Because HPA depends on metrics to know current utilization.
Q95: Why is readiness important to HPA scaling?
Because HPA should only scale based on active and ready pods.
Q96: What is a slow or under-served user experience due to HPA?
It occurs when the workload scales too slowly or not at all under demand.
Q97: What is over-scaling due to HPA?
It occurs when the workload scales too aggressively and increases cost without benefit.
Q98: What is a high target utilization?
It tends to scale later and thus may delay handling bursts.
Q99: What is a low target utilization?
It tends to scale earlier, leading to more replicas and higher cost.
Q100: Why is HPA not a replacement for good capacity planning?
Because autoscaling still needs a solid baseline of app sizing and resource requirements.
Intermediate
Q101: What is a `metrics.k8s.io` API?
It is the standard metrics API for CPU and memory metrics used by HPA.
Q102: What is the Metrics API server deployment?
It is often deployed as `metrics-server` in the cluster.
Q103: Why is metrics-server not always present?
It is an optional cluster component and must be installed.
Q104: What is a custom metrics API?
It is an API that exposes application-specific or external metrics to the HPA.
Q105: What is a custom metrics adapter?
It adapts external metrics to the Kubernetes autoscaling API.
Q106: Why do custom metrics support more realistic scaling?
Because some apps do not scale directly with CPU or memory.
Q107: What is a queue depth metric in HPA?
A queue depth metric can indicate backlog or pending work to scale based on demand.
Q108: What is transaction rate scaling?
Scaling based on application transactions or requests per second.
Q109: Why is API request rate a good HPA signal?
Because it often correlates directly with user demand.
Q110: What is the relationship between HPA and Service load balancers?
The Service load balancer receives traffic, and HPA changes pod count to absorb it.
Q111: Why is ingress traffic relevant to HPA?
Ingress traffic is often a strong driver for scale-out events.
Q112: What is a KEDA?
KEDA (Kubernetes Event-Driven Autoscaling) is a scaler framework for event-driven workloads using custom metrics and queue-based scaling.
Q113: What is KEDA for?
It helps scale workloads based on events, message queues, and external systems.
Q114: Why is KEDA often used with HPA?
Because queue-based or event-driven workloads often do not scale well on CPU alone.
Q115: What is a Prometheus adapter?
It exposes Prometheus metrics to HPA as custom metrics.
Q116: Why is Prometheus common in HPA environments?
Because many clusters already gather Prometheus metrics of application performance and load.
Q117: What is an external metric source?
It is usually a metric not in the cluster itself, like cloud queue length or an API usage counter.
Q118: What is a targetAverageValue?
It is the target value for average custom metrics in an HPA.
Q119: Why are target values often set to a business threshold?
Because they reflect actual working capacity rather than generic resource percentages.
Q120: What is a `currentMetricValue`?
It is the current observed value of the metric used by HPA.
Q121: Why is actual metric value important?
Because the scale decision depends on current vs target.
Q122: What is a desired replica formula roughly?
Desired replicas ≈ current replicas * (current metric / target metric)
Q123: Why does the formula matter?
Because it shows how HPA scales in proportion to observed demand.
Q124: What is a scaling event with stabilization?
The HPA may delay a decision until a period of stable metrics or until the same value persists.
Q125: Why do changes in metrics cause scale churn?
Because if metrics fluctuate too much, HPA might quickly scale up and down.
Q126: What is an HPA overreacting to a short burst?
This occurs when the stabilization window is too short or target thresholds are too aggressive.
Q127: What is underversizing or oversizing a workload?
Undersizing means too few replicas for demand. Oversizing means too many replicas for actual need.
Q128: Why do HPA and resource limits interact?
Because HPA uses resource utilization information, which is affected by pod requests and CPU limits.
Q129: What is a CPU request mismatch?
If the CPU request is too low or too high, HPA may scale wrongly.
Q130: Why does readiness matter to HPA?
Pods that are not ready should not count toward the metric or traffic serving.
Q131: What is a pod readiness metric?
It ensures HPA only uses eligible pods in the scale decision.
Q132: What is a delay before scale-out?
Some clusters or policies intentionally wait to avoid reacting to brief spikes.
Q133: What is a delay before scale-in?
Scale-in is often delayed to avoid churn after a brief drop in demand.
Q134: Why do scale-in policies differ from scale-out?
Because cluster cost and service stability make slow scale-down more cautious.
Q135: What is a scaling decision under multi-metric HPA?
It picks the replica count required by all relevant metrics according to the policy.
Q136: What is a `math` formula in HPA?
It is the calculation used to determine the desired replica count under certain metric types.
Q137: Why is scaling not purely deterministic?
Because metrics, polls, and workloads vary over time, so HPA uses a controller loop.
Q138: What is resource utilization calibration?
It is the tuning process of setting requests and targets to align with actual workload behavior.
Q139: Why do operators add custom metrics for queue depth?
Because queue depth can be a better backlog indicator than CPU.
Q140: What is app throughput scaling?
It uses requests per second or transaction rate as the scaling signal.
Q141: Why is throughput sometimes better than CPU?
Because throughput is directly tied to user demand.
Q142: Why is HPA sometimes paired with VPA?
VPA (Vertical Pod Autoscaler) can adjust pod resource requests and limits while HPA handles replicas.
Q143: What is the difference between HPA and VPA?
HPA changes pod count; VPA changes pod size.
Q144: Why do they sometimes work together?
They handle different dimensions of scaling: count and size.
Q145: What is `HorizontalPodAutoscaler` object definition?
It includes scale target, metrics, min/max replicas, and behavior.
Q146: Why are HPA objects declarative?
Because they describe desired scaling behavior in YAML and are reconciled by the controller.
Q147: Why do cluster operators care about scaling policy?
Because it affects cloud cost, latency, and cluster stability.
Q148: What is a scaling threshold tuning failure?
It happens when thresholds are set too low or too high and do not match workload behavior.
Q149: How do you debug HPA issues?
Check metrics-server, resource requests, target values, pod readiness, and custom metrics.
Q150: What is a misconfigured HPA?
It may never scale, scale too aggressively, or scale based on the wrong metric.
Q151: Why should target metrics be realistic?
Because unrealistic metrics create noisy and unstable scaling decisions.
Q152: What is an application metric like `requestspersecond`?
It directly reflects user demand and can be more meaningful than CPU.
Q153: What is `maxReplicas` used for?
It prevents runaway cost or over-scaling under heavy load.
Q154: What is `minReplicas` used for?
It keeps a baseline number of replicas available even during low traffic.
Q155: Why does low traffic often still need a baseline?
Some workloads require a minimum footprint for responsiveness or warm-up.
Q156: What is a no-scaling condition?
It occurs when metrics are stable and within the target bands.
Q157: Why do some clusters disable HPA by default?
Because some workloads are not well-suited for autoscaling or need custom metrics and tuning.
Q158: What is capacity headroom?
It is spare capacity beyond current load that allows good scaling behavior and latency.
Q159: What can happen during scaling up?
New pods take time to start, so scaling may lag behind traffic.
Q160: Why is startup time important for HPA?
Because HPA must account for the time to create new replicas before demand spikes too much.
Q161: What is a bursty workload?
A workload with sudden spikes in traffic or processing demand.
Q162: Why does HPA need a stabilization window?
Because short-lived bursts should not immediately trigger large or oscillating scale changes.
Q163: What is a scale event and how is it observed?
It is visible via HPA status and the workload replica count.
Q164: Why do custom metrics often require adapters?
Because Kubernetes core metrics do not cover application-specific business metrics.
Q165: What is scale-to-zero?
Some serverless or event-driven systems can scale to zero when idle.
Q166: Why is scale-to-zero not universal for Kubernetes?
Because many workloads must remain warm or ready to respond quickly.
Q167: What is a cold start?
A cold start is the delay before a new replica becomes ready to handle traffic.
Q168: Why do cold starts matter?
Because they affect user latency and the speed needed for HPA to react.
Q169: What is a user-facing latency problem due to HPA lag?
It occurs when traffic increases faster than pods can scale.
Q170: Why do target thresholds need to reflect real traffic?
Because otherwise autoscaling may be too slow or too aggressive.
Q171: What is a queue backlog?
It is the amount of work waiting to be processed.
Q172: Why is queue backlog a strong HPA metric?
Because it can often predict service saturation even when CPU is not high.
Q173: What is SLO?
SLO (Service Level Objective) is a target for acceptable application performance or reliability.
Q174: Why tie HPA to SLOs?
Because scaling decisions should preserve service quality, not just maximize capacity.
Q175: What is HPA-driven cost optimization?
It reduces idle spend by scaling down during low traffic, while satisfying SLOs.
Q176: What is HPA for background jobs?
It can scale workers based on queue depth or processing backlog.
Q177: What is a threshold for background job queue depth?
It determines when additional workers should be created.
Q178: Why are HPA decisions often tuned for specific app types?
Because web traffic, async workers, and databases behave differently.
Q179: Why are application-level metrics more useful than raw CPU?
Because CPU is a generic measure, not necessarily the real bottleneck.
Q180: What is the goal of HPA?
To keep the workload close to the desired demand and maintain service quality while minimizing wasted capacity.
Advanced / Expert
Q181: What is the HPA control loop?
It is the periodic controller reconciliation loop that compares metrics to targets and updates replica counts.
Q182: Why is the HPA control loop important?
Because it defines the cadence and responsiveness of scaling behavior.
Q183: What is a scale-out step policy?
It is a policy setting how much to increase replicas per scaling event.
Q184: What is a scale-down step policy?
It determines how many replicas to remove per scale-down event.
Q185: Why are step policies useful?
They prevent sudden large swings or overreactive adjustments.
Q186: What is a `stabilizationWindowSeconds`?
It is the amount of time HPA waits before recognizing a new scaling trend.
Q187: Why are stabilization windows important in dynamic workloads?
Because metrics can fluctuate and produce noisy scale decisions.
Q188: What is a policy `changePercent`?
It defines the percent change allowed per scaling event.
Q189: How is custom scaling computed?
It depends on the custom metric target and the current metric value for the workload.
Q190: Why are custom metrics important for stateful or event-driven systems?
Because CPU is not always a direct signal of demand or backlog.
Q191: What is KEDA’s architectural relationship with HPA?
KEDA extends HPA by providing scaling based on external or event-driven metrics.
Q192: Why is KEDA often preferred for Kafka or queue-based systems?
Because such systems often scale according to queue depth or backlog length more than CPU.
Q193: What is an external metric adapter?
It bridges a third-party metric source to the HPA API.
Q194: Why is metric adapter reliability critical?
Because an unstable metric source can lead to wrong or dangerous scale decisions.
Q195: What is a metric-source failure?
It happens when the metrics adapter or backend cannot provide valid data to HPA.
Q196: Why does HPA need healthy metrics?
Because without them, it cannot compute a correct replica count.
Q197: What is the relation between HPA and SLOs?
HPA should be tuned to keep service-level objectives within target thresholds while optimizing cost.
Q198: What is a scale-out safety check?
It is a policy or condition ensuring scaling does not exceed safe or cost-effective limits.
Q199: What is a scale-down safety check?
It ensures scaling down does not reduce capacity below acceptable service levels.
Q200: What is the main lesson of Kubernetes HPA?
HPA is the cluster’s automatic capacity controller: it measures demand, compares it to policy and target values, and adjusts replica counts to balance performance, cost, and resilience.