Kubernetes Resource Limits

Kubernetes Resource Limits


Beginner

Q1: What are Kubernetes resource limits?

Resource limits are constraints placed on CPU and memory usage for containers or Pods.

Q2: Why do we need resource limits?

To avoid a single pod consuming too much cluster or node resources and degrading other workloads.

Q3: What is CPU in Kubernetes?

CPU is the compute processing capacity available to a container or Pod.

Q4: What is memory in Kubernetes?

Memory is the RAM available to a container or Pod.

Q5: What is a request in Kubernetes?

A request is the minimum amount of CPU or memory reserved for a container.

Q6: What is a limit in Kubernetes?

A limit is the maximum amount of CPU or memory a container can use.

Q7: What is the difference between request and limit?

Request is guaranteed allocation; limit is the maximum cap.

Q8: Why do requests matter for scheduling?

The scheduler uses requests to decide which node can run the Pod.

Q9: Why do limits matter for protection?

Limits prevent runaway processes from exhausting node resources.

Q10: What is the `resources` block in a Pod spec?

The `resources` block contains `requests` and `limits` for CPU and memory.

Q11: What is CPU request?

CPU request is the amount of CPU the scheduler tries to reserve for the Pod.

Q12: What is CPU limit?

CPU limit is the upper bound on CPU the container can consume.

Q13: What is memory request?

Memory request is the amount of memory guaranteed to the Pod by the scheduler.

Q14: What is memory limit?

Memory limit is the maximum memory the container can use before it is killed or throttled.

Q15: What happens when a container exceeds memory limit?

Usually the container is killed by the kernel OOM killer or the runtime enforces the limit.

Q16: What happens when a container exceeds CPU limit?

The container may be throttled or slowed, but it is not usually killed immediately.

Q17: What is QoS in Kubernetes?

QoS (Quality of Service) classifies Pods based on request and limit values to determine eviction priority.

Q18: What are the Kubernetes QoS classes?

  • Guaranteed
  • Burstable
  • BestEffort

Q19: What is Guaranteed QoS?

Guaranteed QoS means requests and limits are equal for all containers.

Q20: What is Burstable QoS?

Burstable QoS means requests and limits are set, but not equal.

Q21: What is BestEffort QoS?

BestEffort means no requests or limits are set.

Q22: Why are QoS classes important?

They affect eviction priority and cluster behavior during resource pressure.

Q23: What is node pressure?

Node pressure means the node is short on compute, memory, or disk and Kubernetes starts evicting workloads.

Q24: Why are limits important for cluster stability?

Because without limits, one workload can starve others or trigger node instability.

Q25: What is CPU throttling?

CPU throttling slows a process when it exceeds the assigned CPU quota.

Q26: What is OOM kill?

OOM kill means the kernel kills a process because it exceeded memory limits or system memory pressure.

Q27: Why do memory limits often need careful tuning?

Because memory limit too low causes crashes; too high can starve other pods.

Q28: What is a request-to-limit ratio?

It is the relationship between reserved and allowed resources, affecting burst capacity.

Q29: What is the scheduler?

The scheduler decides which node can run a Pod based on resource requests and the cluster state.

Q30: What is node allocatable capacity?

It is the CPU and memory available for scheduling after system overhead and reservation.

Q31: Why does the scheduler care about requests?

Because the Pod can only schedule if the node has enough resource requests for the workload.

Q32: Why might a pod stay Pending?

A pod can stay pending if no node has enough available resources matching its requests.

Q33: What is a Pending pod?

A Pending pod is waiting to be scheduled or bound to a node.

Q34: What is `cpu: "500m"`?

500m means 0.5 CPU (half a CPU core).

Q35: What is `memory: "512Mi"`?

512Mi means 512 mebibytes of memory.

Q36: What is the difference between Ki, Mi, Gi, and m values in Kubernetes?

They are units for memory and CPU, where m means millicpu and Ki/Mi/Gi are binary memory units.

Q37: Why is CPU usually measured in millicores?

Because CPU is often fractional and expressed as thousandths of a core.

Q38: What is a resource unit for memory?

Common memory units include bytes, Ki, Mi, Gi, and similar binary units.

Q39: Why should workloads set requests?

Because without requests, the scheduler may place workloads unpredictably and overload nodes.

Q40: Why should workloads set limits?

Because it enforces a ceiling and reduces the chance of noisy neighbor behavior.

Q41: What is a noisy neighbor?

A noisy neighbor is a workload using excessive resources and impacting other workloads on the same node.

Q42: What is a node-level resource pressure event?

It is when the node is running out of CPU, memory, or disk and Kubernetes starts corrective actions.

Q43: What is eviction?

Eviction is the process Kubernetes uses to terminate pods when the node can no longer support them.

Q44: Why do evictions matter?

Evictions can disrupt workloads if their resource requests and limits are poorly configured.

Q45: What is a resource quota?

A ResourceQuota limits the total amount of resources a namespace can consume.

Q46: Why is a ResourceQuota useful?

It prevents a single namespace from consuming too much cluster capacity.

Q47: What is a LimitRange?

A LimitRange defines default or maximum resource values for pods in a namespace.

Q48: Why use a LimitRange?

To enforce a baseline and default resources for workloads in a namespace.

Q49: What is a default request or default limit?

Kubernetes can apply defaults to pods that omit resource values based on the LimitRange.

Q50: What is a resource request for a Deployment?

A Deployment’s pod template can include resource requests and limits for all pod replicas.

Q51: How do limits affect autoscaling?

Autoscaling decisions may depend on observed pod resource utilization and requests.

Q52: What is HPA?

HPA (Horizontal Pod Autoscaler) scales pods based on metrics such as CPU utilization.

Q53: Why are correct limits important for HPA?

Because HPA uses resource utilization to decide when to scale, and bad limits can distort the metrics.

Q54: What is a CPU utilization metric?

It is the average CPU percentage used by the pod relative to its request or limit.

Q55: Why is utilization not the same as absolute usage?

Because the same percentage may correspond to different load depending on the pod’s request size.

Q56: What is a memory usage metric?

It is the actual memory measured by the pod or node.

Q57: What is a container crash due to OOM?

It happens when the process attempts to use more memory than allowed and is killed.

Q58: What is `resources: {}` in a pod spec?

It means no explicit requests or limits are set for the container.

Q59: Why is setting nothing risky?

Because the pod may be scheduled without guarantees and may starve or be starved by others.

Q60: Why should limits be realistic?

Because unrealistic limits create throttling, crashes, or overprovisioning.

Q61: What is a memory leak?

A memory leak is a bug causing a program to consume memory continuously without releasing it.

Q62: Why do memory limits help with memory leaks?

They cap damage and can trigger restarts when a workload becomes unhealthy.

Q63: What is a CPU hog?

A CPU hog is a process or workload consuming excessive CPU time.

Q64: Why do CPU limits matter for CPU hogs?

They prevent a single workload from monopolizing all compute on a node.

Q65: What is resource pressure in a cluster?

It is when the cluster experiences shortage relative to requests and usage.

Q66: What is scheduler placement?

It is the decision of which node should run a pod based on resources, labels, and constraints.

Q67: What is a node selector?

A node selector constrains which nodes a pod may be scheduled onto.

Q68: Why are node selectors relevant to resource limits?

Because workloads may need to be placed onto nodes with the right capacity or hardware.

Q69: What is a taint and toleration?

Taints and tolerations allow pods to schedule onto specific nodes or avoid others.

Q70: Why do taints and tolerations affect resources?

Because they influence where workloads run and how resources are shared.

Q71: What is a pod disruption budget?

It limits how many pods can be disrupted during voluntary or involuntary disruptions.

Q72: Why are resource limits tied to resilience?

Because if a node is overloaded, more pods may be evicted and disruption increases.

Q73: What is a resource threshold?

A threshold is a limit or alerting point that flags unusual resource consumption.

Q74: What is a request vs limit best practice?

Use requests for scheduling and guarantees, and limits for safe caps based on real workload needs.

Q75: What is a memory request for a Java app?

It often includes heap and runtime overhead plus some buffer for GC and OS memory.

Q76: Why is Java memory tuning tricky?

Java uses heap, object overhead, native memory, and GC behavior that affect actual memory consumption.

Q77: Why are CPU limits sometimes not suitable for all workloads?

Some workloads benefit from bursting or using idle CPU when available.

Q78: What is a CPU burst?

A CPU burst is temporary additional CPU usage above the requested amount, if allowed by the node.

Q79: Why do strict limits lower burst capacity?

Because limits cap the maximum average throughput a process can use.

Q80: What is a kill due to OOM in Kubernetes?

A pod can be restarted when the container exceeds its memory limit and is terminated.

Q81: What is a memory limit plus cgroup behavior?

Kubernetes uses Linux cgroups to enforce memory limits.

Q82: Why do Linux cgroups matter?

They are the underlying mechanism used by Kubernetes to enforce limits.

Q83: What is `cgroup` memory enforcement?

It restricts memory usage to the configured limit and may kill the process when exceeded.

Q84: Why should you monitor node usage?

Because a cluster can look healthy until one node saturates and causes evictions.

Q85: What is a node starvation condition?

It is when a node does not have enough resources left to satisfy workloads.

Q86: What is a resource leak?

A resource leak is an application bug that gradually consumes resources.

Q87: Why do resource limits not fix all app bugs?

They can mask failures by causing restarts rather than correcting the root cause.

Q88: What is a request or limit mismatch?

It happens when the numbers do not reflect actual runtime behavior, causing poor scheduling or failures.

Q89: What is a noisy neighbor in a Kubernetes cluster?

A workload with misconfigured or greedy limits can starve others on the same node.

Q90: What is capacity planning?

Capacity planning estimates how much CPU and memory the cluster must provide for current and future workloads.

Q91: What is a resource reservation?

A reservation is a guarantee that some resources are left for system daemons or critical workloads.

Q92: Why reserve resources for the system?

Because system services also need CPU and memory, and they should not be starved.

Q93: What is a resource request for the kubelet or system daemons?

These system-level resources are important for node health and should not be fully consumed by user workloads.

Q94: What is cluster overcommit?

Overcommit happens when total requested resources exceed total cluster capacity.

Q95: Why is cluster overcommit risky?

It can lead to node pressure and pod eviction under load.

Q96: Why is pod sizing important?

Poor pod sizing causes poor utilization, unexpected throttling, or crash loops.

Q97: What is a resource utilization baseline?

It is the typical amount of CPU and memory a workload uses in production.

Q98: What is performance tuning for resources?

It means tuning requests and limits based on real load and throughput observations.

Q99: What is autoscaling vs manual resource tuning?

Autoscaling responds to traffic dynamics; tuning sets a baseline for pod sizing and resource allocation.

Q100: What are common resource misconfigurations?

  • no requests
  • no limits
  • requests too low
  • limits too low
  • limits too high
  • ignoring node capacity

Intermediate

Q101: What is a `requests` value used by scheduler?

The scheduler uses requests to find a suitable node and reserve space.

Q102: What is a `limits` value used by runtime?

The runtime uses limits to enforce maximum cgroup usage.

Q103: Why do requests and limits differ in practice?

Because the runtime and scheduler use them differently: one is soft reserved capacity, the other is enforcement.

Q104: Why do requests matter more than limits for scheduling?

Because the scheduler can only place pods on nodes with enough available requests.

Q105: Why are requests sometimes set equal to limits?

When you want predictable behavior and consistent capacity planning.

Q106: Why are requests sometimes lower than limits?

To allow burst capacity without starving the node.

Q107: What is the CPU burst model?

It allows short spikes above the request when there is available node capacity.

Q108: What is a CPU throttle threshold?

It is the point at which the runtime limits or throttles CPU usage.

Q109: Why is memory overcommit dangerous?

Because it can cause OOM kills and disruption when multiple pods compete for memory.

Q110: Why can high memory limits cause poor utilization?

Because the cluster may reserve memory that is not actually needed, reducing scheduling density.

Q111: What is `EmptyDir` and resource use?

It is local ephemeral storage, not CPU or memory, but may still affect disk usage.

Q112: What is the resource use of system daemons?

The kubelet, kernel, and control plane components also consume CPU and memory.

Q113: What is node allocatable memory?

The actual memory left for application workloads after system reservations.

Q114: What is a pressure condition on a node?

It occurs when memory or disk is under heavy strain and Kubernetes starts eviction or scheduling restrictions.

Q115: What is a `PriorityClass`?

A PriorityClass gives a pod higher or lower scheduling and eviction priority.

Q116: Why do priorities matter with resource limits?

Because under pressure, Kubernetes may evict lower-priority pods first.

Q117: Why are cluster-level quotas useful?

They help control the total resources users or namespaces can consume.

Q118: What is a quota object?

A quota object defines total requests and limits for a namespace.

Q119: Why use `LimitRange` plus quotas together?

Because LimitRange sets defaults and bounds, while quota limits total usage.

Q120: What is a resource-level namespace boundary?

It is the namespace-level total for CPU, memory, and storage.

Q121: What is a cluster autoscaler?

The cluster autoscaler adds or removes nodes based on pending pods and cluster load.

Q122: Why is autoscaler relevant to resource limits?

Because I/O and CPU pressure dictate whether more nodes are needed.

Q123: What is a node pool?

A node pool is a group of nodes with similar size or characteristics, often for scheduling CPU/memory workloads.

Q124: Why do some clusters create node pools for different workload types?

Because different apps have different CPU, memory, and capacity needs.

Q125: What is a memory request vs memory usage metric?

Request is the reserved baseline; usage is actual observed allocation.

Q126: Why is usage monitoring important?

Because pods can exceed requests or limits without the platform realizing it.

Q127: What is a dashboard or metrics server?

It shows pod resource usage over time and helps tune requests and limits.

Q128: What is a CPU saturation event?

A saturation event happens when the node or pod is at or near its CPU capacity.

Q129: What is memory saturation?

It is when memory usage is close to limit or near the node’s total capacity.

Q130: Why are resource requests important for cluster capacity?

They are the basis for scheduling and future node scaling.

Q131: Why are limits sometimes omitted for sidecars?

Because sidecars may not need a strict cap or may be designed to burst.

Q132: What is a sidecar container?

A sidecar is a helper container in the same pod that shares resources with the main app.

Q133: Why are sidecars relevant to resource budgeting?

Because sidecars compete for the same pod-level resource pool.

Q134: What is a pod-level resource sum?

The sum of all containers’ requests and limits in a pod.

Q135: Why is pod-level sum relevant?

Because pod scheduling and resource enforcement happen at the pod/container aggregate level.

Q136: What is a memory request for a sidecar?

It may be lower than the main app, but the total still matters.

Q137: Why should service and sidecar resources be considered together?

Because they share the pod’s total and can affect each other.

Q138: What is a “too low limit” problem?

It can cause the app to be throttled or killed even though the node has enough overall capacity.

Q139: What is a “too high limit” problem?

It can starve other pods or reduce cluster density, leading to poor scheduling.

Q140: Why does resource tuning require experimentation?

Because real workloads differ by runtime, concurrency, and traffic patterns.

Q141: What is a container OOM with restart policy?

A container can restart after OOM, which can cause a crash loop if limits are set too low.

Q142: Why do well-sized limits reduce restarts?

Because memory and CPU constraints remain aligned with actual workload behavior.

Q143: What is a memory headroom?

Headroom is the extra memory left after requests and limits to absorb bursts and system activity.

Q144: Why is headroom important?

It helps avoid unnecessary evictions and OOM kills.

Q145: What is a CPU headroom?

The extra CPU capacity in the node that allows bursts without system starvation.

Q146: Why are node-level limits smaller than total node capacity?

Because operating system daemons and Kubernetes components also need resources.

Q147: What is a scheduling score?

Kubernetes uses scoring to pick the best node based on many factors, including resource fit.

Q148: What are generic resource names?

CPU and memory are the standard ones, but custom resources can also be used in some clusters.

Q149: What is a custom resource for GPU?

GPUs can be requested and limited as custom resources in some clusters.

Q150: What is a GPU resource limit?

It caps how much accelerator capacity a pod can use and is often modeled separately from CPU and memory.

Q151: What are node resource reservations?

They keep some CPU and memory available for system-critical tasks.

Q152: Why is node reservation important?

Because the node still needs to run kubelet, container runtime, and OS tasks.

Q153: What is a service-level resource plan?

It is the policy and sizing model for each production service based on expected traffic and workload characteristics.

Q154: Why is baseline usage important?

Because it informs safe requests and limits, which impact scheduling and reliability.

Q155: What is cluster capacity planning?

It is the process of understanding total available CPU and memory and how much is consumed by workloads.

Q156: What is load testing and resource measurement?

It is the process of measuring a workload under realistic conditions to determine good limits.

Q157: What is a canary resource test?

It measures whether a new workload version fits its resource budget before full rollout.

Q158: What is resource observability?

It is the visibility into actual CPU, memory, and process behavior of workloads.

Q159: What is a resource alert?

An alert signals abnormal usage above expected thresholds.

Q160: What is a resource usage trend?

A trend shows whether a workload grows over time and may require tuning or scaling.

Q161: Why is resource tuning an iterative process?

Because workloads evolve and traffic changes over time.

Q162: What is a pod left with no requests and limits?

It receives no scheduling guarantee and can be noisy or starved on a node.

Q163: Why is overprovisioning dangerous?

It reduces scheduling efficiency and can create unanticipated contention.

Q164: What is garbage collection and resources?

Garbage collection may generate temporary CPU and memory use that is not always visible in baseline app metrics.

Q165: What is the relation between resource requests and evictions?

Under resource pressure, pods with lower QoS or lower priority are more likely to be evicted.

Q166: What is a resource pressure event in a cluster?

It is when the cluster experiences insufficient CPU or memory relative to current demand.

Q167: Why is QoS a key concept?

Because it affects eviction order and resilience under stress.

Q168: What is the bounded usage model?

It is the idea that workloads use only the resources they are allowed to use based on limits and requests.

Q169: What is a slow or starved pod?

A pod starved of CPU or memory can behave erratically or fail under load.

Q170: What is memory pressure?

Memory pressure increases the risk of OOM kills and performance degradation.

Q171: What is CPU pressure?

CPU pressure makes the process slow or queued behind resource limits.

Q172: Why do both CPU and memory matter?

Because one can cause slowness while the other can cause failures or restarts.

Q173: What is a pod quality-of-service policy?

It determines which pods are sacrificed first during node pressure events.

Q174: Why is resource-aware scheduling crucial?

Because it ensures workloads land on suitable nodes and avoids chaotic contention.

Q175: Why should operators prefer explicit requests and limits?

It improves reliability, predictability, and cluster health.

Q176: What is a baseline setting for app resources?

It is the initial request and limit values set based on profiling and expected load.

Q177: What is cluster-level resource fairness?

It ensures users or namespaces do not overwhelm the shared cluster resources.

Q178: What is request aggregation?

It is the total requests across all pods in a namespace or cluster.

Q179: What is a resource utilization ratio?

It is the actual usage relative to request or limit values.

Q180: Why is measuring real usage crucial for resource tuning?

Because good resource limits are derived from actual workload behavior, not guesses.

Advanced / Expert

Q181: What is cgroup v1 vs cgroup v2?

They are different Linux cgroup versions used by the runtime to enforce resource limits.

Q182: Why is cgroup version relevant?

Because some tools and kernels behave differently depending on the cgroup version.

Q183: What is resource accounting?

It is the tracking of CPU and memory consumed by containers and processes.

Q184: Why is accounting needed?

Because it provides the raw input for scheduling, alerting, autoscaling, and tuning.

Q185: What is `top` or `kubectl top`?

They show current resource usage of nodes or pods for operational observability.

Q186: Why is `kubectl top` valuable?

It helps estimate whether requests and limits are realistic and whether workloads are saturating the cluster.

Q187: What is kernel OOM killer behavior?

The OOM killer may kill the highest-memory container or process when memory is exhausted.

Q188: Why does OOM kill matter in Kubernetes?

Because it can restart a pod abruptly and cause service disruption.

Q189: What is a memory burst beyond request?

Pods may temporarily use more memory than their request if there is free node memory.

Q190: Why is burst possible but limited?

Because the node can allow short spikes if capacity remains available.

Q191: What is a scheduler fit problem?

It happens when a pod cannot be scheduled because no node has enough requested resources.

Q192: Why is careful CPU request budgeting important?

Because mis-sized CPU requests can cause under-scheduling or overutilization.

Q193: What is a high-concurrency app?

An app with many parallel workers or requests that may use bursts of CPU and memory.

Q194: Why do bursty apps need careful limits?

Because they can be starved or killed if limits are too low and not aligned with traffic spikes.

Q195: What is a latency-sensitive workload?

It depends on low response time and stable resource availability.

Q196: Why do latency-sensitive apps need better resource reservations?

Because if they do not have enough resources, they may queue or throttle and increase latency.

Q197: What is a low-latency memory profile?

It often requires more predictability and less aggressive overcommit.

Q198: Why is CPU limit too low a common production issue?

Because apps can degrade under load and appear “slow” even though they are still running.

Q199: What is the relationship between resource limits and autoscaling?

Autoscaling decides scale-out based on load, while limits decide how much each pod can consume before throttling or OOM.

Q200: What is the main lesson of Kubernetes resource limits?

Resource limits are about safety, scheduling correctness, and cluster health; requests reserve capacity, limits cap usage, and both must be based on real workload behavior to keep the cluster stable and the app reliable.