Kubernetes Probes
Kubernetes Probes
Beginner
Q1: What is a Kubernetes probe?
A Kubernetes probe is a health check that tells the kubelet whether a container is alive, healthy, and ready to receive traffic.
Q2: Why do we need probes?
Probes help Kubernetes decide when to restart a container, when to route traffic to it, and whether it is healthy.
Q3: What is the kubelet?
The kubelet is the node agent that runs inside each node and manages the containers in Pods.
Q4: What is a container lifecycle?
The container lifecycle includes startup, running, readiness, and termination.
Q5: What is a readiness probe?
A readiness probe checks whether a container is ready to accept traffic.
Q6: What is a liveness probe?
A liveness probe checks whether a container is alive and should be restarted if it is unhealthy.
Q7: What is a startup probe?
A startup probe checks whether the application has completed startup before liveness and readiness checks begin.
Q8: Why are readiness probes important?
They prevent traffic from reaching containers that are not yet ready.
Q9: Why are liveness probes important?
They allow Kubernetes to restart containers that have become stuck or unhealthy.
Q10: Why are startup probes important?
They help slow-starting applications avoid being restarted prematurely by liveness probes.
Q11: What is a probe action?
A probe action is how Kubernetes checks health, such as HTTP, TCP, or command execution.
Q12: What is an HTTP probe?
An HTTP probe sends an HTTP request to a specific path and expects a status code response.
Q13: What is a TCP probe?
A TCP probe checks whether a port is open and accepting connections.
Q14: What is an exec probe?
An exec probe runs a command inside the container and succeeds or fails based on exit code.
Q15: What is the probe failure threshold?
The threshold is the number of consecutive failures before a probe is considered failed.
Q16: What is probe period?
The period is how often the kubelet runs the probe.
Q17: What is probe timeout?
The timeout is how long the probe waits before considering the check failed.
Q18: What is initial delay?
Initial delay is how long the kubelet waits before starting the first probe.
Q19: What is failure threshold?
Failure threshold is how many consecutive failures are required before a probe transitions to failure.
Q20: What is success threshold?
Success threshold is how many consecutive successful checks are required before a probe is considered successful again.
Q21: Why does a readiness probe matter for Services?
Because Services only route traffic to ready Pods.
Q22: What is a Pod condition?
A Pod condition reflects the current state of readiness, liveness, or startup.
Q23: What is a ready condition?
A Pod is marked ready when all readiness probes pass and the pod is ready to receive traffic.
Q24: What does it mean if a Pod is not ready?
It means Kubernetes will not send traffic to it via a Service unless special routing rules apply.
Q25: What happens when a liveness probe fails?
The container is restarted by the kubelet unless the container is not restartable.
Q26: Why should liveness probes be careful?
A liveness probe that is too aggressive can restart healthy containers unnecessarily.
Q27: What is a probe target path?
In an HTTP probe, the target path is the URL path to request, such as /healthz or /ready.
Q28: What is a probe port?
The probe port is the port the kubelet uses to connect for HTTP or TCP probes.
Q29: Why use /healthz or /ready?
These are conventional endpoints often implemented by apps to report internal health.
Q30: What is app health?
App health is whether the application is functioning correctly and able to serve traffic.
Q31: What is a process health check?
A process health check checks whether the main process is alive and responding.
Q32: What is a dependency check?
A dependency check verifies that an app can reach required databases, caches, or other services.
Q33: Why should readiness include dependencies?
If the app cannot reach a database, it may not be ready even though the process is running.
Q34: What is a startup check for slow boot?
It ensures the app has enough time to initialize before liveness begins.
Q35: Why do some workloads need startup probes?
Because their initialization can take minutes, and liveness would otherwise restart them.
Q36: What is a probe configuration example?
Example:
- readinessProbe: HTTP GET on /ready
- livenessProbe: TCP port 8080
- startupProbe: HTTP GET on /startup
Q37: What is a rapid restart loop?
A rapid restart loop occurs when a container keeps failing health checks and being restarted.
Q38: Why is a rapid restart loop dangerous?
It can create instability and hide the underlying problem.
Q39: What is a production probe?
A production probe is a health check designed for live traffic and operational reliability.
Q40: Why can probes be wrong?
Because a probe may test the wrong thing or use logic that doesn’t reflect actual readiness.
Q41: What is a false positive probe?
A false positive means the probe says healthy when the application is actually broken.
Q42: What is a false negative probe?
A false negative means the probe says unhealthy when the application is actually working.
Q43: Why must probes be accurate?
Because they affect traffic routing and restarts.
Q44: What is a health endpoint?
A health endpoint is an HTTP URL that returns status to indicate app health.
Q45: What is a readiness endpoint?
A readiness endpoint indicates the application is ready to accept real traffic.
Q46: What is a liveness endpoint?
A liveness endpoint indicates the process is alive, even if it is not ready for traffic.
Q47: Why separate readiness and liveness?
Because being alive is not the same as being ready to serve requests.
Q48: What is a slow-starting app?
An app whose startup takes longer than a few seconds and needs delayed liveness checks.
Q49: What is a long-running job?
A long-running job may not be a web service and may not need normal readiness semantics.
Q50: What is an init container?
An init container runs before app containers start and can perform setup or dependency checks.
Q51: Why do init containers sometimes help with startup probes?
They can do one-time setup before the main app starts.
Q52: What is readiness gating?
Readiness gating prevents traffic from reaching a pod until it is ready.
Q53: What is a Pod being ready?
It means the app has passed readiness checks and is included in Service endpoints.
Q54: What is a readiness check for dependencies?
It may test DB connectivity, cache connectivity, or external service availability.
Q55: What is an HTTP status code threshold?
A probe may expect 200-399 as success, and other codes as failure.
Q56: What is a TCP port readiness check?
It only checks socket availability, not application correctness.
Q57: Why is TCP-only readiness sometimes insufficient?
Because the app could accept connections but still fail deeper logic or user requests.
Q58: What is command-based readiness?
It runs a shell command inside the container and expects a zero exit code.
Q59: Why are exec probes sometimes problematic?
They use shell execution inside the container, which adds complexity and performance overhead.
Q60: What is the kubelet probe loop?
It is the periodic sequence in which the kubelet runs configured probes.
Q61: What is a probe serialization?
It is the sequential execution of health checks on the node.
Q62: What is a probe lifecycle event?
It is the moment a container transitions through startup, ready, and alive states according to checks.
Q63: What is a readiness probe vs liveness probe difference?
Readiness controls if traffic is routed to the pod. Liveness controls if the pod should be restarted.
Q64: What does CPU impact from probes look like?
Probes add small overhead but generally are lightweight compared with the application itself.
Q65: Why should probe handlers be efficient?
Because a misconfigured or heavy probe can affect node performance.
Q66: What is a health endpoint that returns 503?
A 503 often indicates the app is not ready or not healthy.
Q67: Why is 200 OK common for readiness?
It clearly signals the app is ready.
Q68: Why use 204 No Content for health endpoints?
It may be a lightweight health response with no payload.
Q69: What is failure tolerance?
Failure tolerance describes how many probe failures are allowed before the app is considered dead or unready.
Q70: What is an app that never becomes ready?
This can happen if the app waits on a dependency or a bad startup path.
Q71: Why is restart loop a symptom of bad probes?
Because the app may be failing the liveness probe repeatedly even though it is not truly stuck.
Q72: What is a startupProbe + readinessProbe combination?
This is often used when the app takes a while to start, but should be handled gracefully.
Q73: What is a Pod not ready event?
It indicates the pod is not included in Service endpoints because readiness checks are failing.
Q74: What is pinging an app with a health endpoint?
It means the probe is checking the app is accepting requests.
Q75: What is a `httpGet` probe?
It is an HTTP probe in Kubernetes manifest syntax.
Q76: What is a `tcpSocket` probe?
It is a TCP socket probe in Kubernetes manifest syntax.
Q77: What is a `exec` probe?
It is a command-based probe in Kubernetes manifest syntax.
Q78: What is a probe port field?
It denotes which container port is used for the probe.
Q79: What is a `host` field in a probe?
Host is not commonly used for probes, but some special wiring may target host-level addresses in advanced scenarios.
Q80: What is a scheme field in HTTP probe?
The scheme selects HTTP or HTTPS for the probe request.
Q81: What is `initialDelaySeconds`?
How long the kubelet waits before performing the first probe check.
Q82: What is `periodSeconds`?
How often the probe runs.
Q83: What is `timeoutSeconds`?
How long the probe waits before considering the check failed.
Q84: What is `failureThreshold`?
The number of consecutive failures before the kubelet marks it unhealthy or unready.
Q85: What is `successThreshold`?
The number of successes required before a failed probe becomes successful.
Q86: Why is kubelet behavior important to understand?
Because the kind of probe and thresholds determine traffic behavior and restarts.
Q87: What is a probe that checks only process liveness?
It can work but often misses application-level problems like database unavailability.
Q88: Why should probes be aligned with app semantics?
Because the app may be alive but still unable to serve user traffic.
Q89: What is a container that exits fast?
A container that exits quickly may not be caught by readiness or liveness if the process dies.
Q90: What is a crash loop?
A crash loop is when a container repeatedly exits and restarts, often due to a failure.
Q91: What is a restart policy?
Restart policy determines whether Kubernetes restarts the container after failure.
Q92: Why do readiness and liveness rely on restart policy?
Because restart policy defines what happens when the container fails.
Q93: What is a pod restart due to liveness?
When the app fails liveness, kubelet restarts the container.
Q94: What is a Pod phase?
A Pod phase includes Pending, Running, Succeeded, Failed, and Unknown.
Q95: Why is readiness distinct from pod phase?
A Pod can be Running but not Ready.
Q96: What is the difference between Running and Ready?
Running means the process exists. Ready means traffic can be routed to it.
Q97: What is traffic routing to ready Pods?
This is the core function of readiness in Kubernetes Services.
Q98: What is ingress and readiness?
Ingress traffic is routed through Services to ready Pods.
Q99: Why do probes matter for external traffic?
Because external systems often rely on the Service and Ingress to route to the correct Pods.
Q100: What is the primary goal of probes?
To ensure the app is alive, ready, and actually able to serve traffic.
Intermediate
Q101: What is the difference between readiness and startup probe semantics?
Readiness is about traffic eligibility. Startup is about not treating a still-starting app as failed too early.
Q102: Why use startupProbe for applications with long startup?
Because liveness could restart the process before it finishes initialization.
Q103: What is a start-up dependency check?
It may check that required databases or caches are reachable before the app becomes ready.
Q104: What is a readiness check for database connection?
It checks that the app can connect to the DB before accepting requests.
Q105: What is a readiness check for external dependencies?
It verifies the app can reach required external services or APIs.
Q106: What is an HTTP readiness endpoint that checks DB connection?
It might return 503 if DB connectivity is broken even though the app process is running.
Q107: Why can liveness be based on process health alone?
It is often enough for CPU-bound processes that either run or not.
Q108: Why can liveness be too aggressive for network-heavy apps?
Network or dependency outages may not be the app’s fault but can make it appear unhealthy.
Q109: What is a probe for slow app initialization?
A startupProbe can give the app enough time before liveness becomes active.
Q110: Why do some apps use /healthz and /ready endpoints differently?
Because /healthz often means “process alive”; /ready means “can serve traffic”.
Q111: What is a custom health status body?
Apps can return JSON or plain text describing health to help debugging.
Q112: Why are probe endpoints often small and cheap?
Because they are called frequently by kubelet and should not be expensive.
Q113: What is a health endpoint with dependencies?
It can combine DB ping, cache ping, and response readiness checks.
Q114: Why avoid heavy dependency checks in readiness?
Because they may become expensive or cause cascading failures if dependencies are slow.
Q115: What is a shallow readiness check?
A shallow readiness check confirms the process is ready to accept traffic without doing deep dependencies.
Q116: What is a deep readiness check?
A deep readiness check validates the app can talk to critical dependencies and is fully functional.
Q117: Why must readiness logic match business expectations?
Because a pod should only receive traffic when the app is truly usable.
Q118: What is a liveness case where the app is deadlocked?
A liveness probe can detect this by failing repeatedly and restarting the container.
Q119: Why does deadlock sometimes require liveness?
Because the process may appear alive but is hung.
Q120: What is a probe that runs against a local port?
This is a common pattern for apps exposing management or health endpoints locally.
Q121: What is the kubelet probe and service behavior interplay?
The kubelet updates pod readiness, which then updates Service endpoints.
Q122: What does readiness influence in Kubernetes?
It influences inclusion in Service endpoints and external routing.
Q123: What is a PodDefault readiness state?
The pod is not ready until readiness checks pass.
Q124: What is a readiness false positive risk?
The app might look ready but not actually be able to serve all traffic.
Q125: What is a readiness false negative risk?
The app might be healthy but kept out of Service traffic due to a probe bug.
Q126: What is an HTTP probe expecting 2xx?
Many health checks treat 2xx as success.
Q127: What is an HTTP probe expecting a non-2xx response?
Some apps intentionally use 503 or 500 to signal not ready.
Q128: Why is it useful to return 503 during readiness?
Because it tells the kubelet and the load balancer that the app is not ready for traffic.
Q129: What is a TCP probe?
It checks if the port accepts a TCP connection.
Q130: Why is TCP probe less semantic than HTTP?
Because it says “port is open,” not “app is healthy”.
Q131: What is an exec probe with `curl –fail`?
It executes `curl` inside the container and fails the probe if the HTTP call is not successful.
Q132: Why do many apps prefer HTTP probes?
Because HTTP semantics map well to app health and readiness.
Q133: Why avoid exec probes when possible?
Because they are heavier and more complex to manage within containers.
Q134: What is a liveness check with `cat /tmp/healthy`?
It is an exec-based check for a file signifying health.
Q135: Why use file-based health checks?
For some applications, a health file or lock file indicates whether to continue serving.
Q136: Why are probes affected by initialization time?
Because the container may need time to warm caches or connect to external services.
Q137: Why are delays configurable?
Because startup time varies across deployments and frameworks.
Q138: What is a startup probe example in manifests?
Example pattern:
- startupProbe: httpGet /startup
- livenessProbe: httpGet /healthz
- readinessProbe: httpGet /ready
Q139: Why is ordering subtle in manifests?
Startup must start first, then readiness/liveness become active after startup completes.
Q140: Why is the lack of a startupProbe a common cause of restart loops?
Because a slow app may fail the liveness probe before it has fully initialized.
Q141: What is a default probe failure behavior?
Once a failure threshold is reached, kubelet marks the container unhealthy or unready.
Q142: What is a readiness probe on Pod termination?
A readiness probe may show failures during teardown to stop new traffic before shutdown.
Q143: Why do operators care about readiness state?
It is a direct signal of whether production traffic should be routed to a pod.
Q144: What are common probe anti-patterns?
- using liveness as readiness
- no probes at all
- overly aggressive thresholds
- testing the wrong dependency
- using heavy checks in every loop
Q145: Why is “liveness = readiness” a common mistake?
Because a pod can be alive but not able to serve traffic; restarting it is not always the solution.
Q146: Why use a dedicated /ready endpoint?
Because it can check a narrower set of conditions than /healthz.
Q147: Why use a dedicated /live endpoint?
Because it can check just whether the process is alive and not in a hung or broken state.
Q148: What is a startup endpoint used for startup probe?
It can confirm the app finished initialization, such as migrations or warm-up.
Q149: What is a probe and ingress interplay?
Ingress and service routing rely on readiness to avoid sending traffic to unready pods.
Q150: What is a readiness outage?
It means a pod is running but not receiving traffic because readiness checks fail.
Q151: Why is a pod with failing readiness bad for service availability?
Because it is effectively removed from the routing pool even though the process may still be alive.
Q152: What is a health endpoint that returns empty or 200?
This is often used for a very simple probe.
Q153: What are common frameworks with built-in health endpoints?
Many frameworks expose /health, /healthz, /actuator/health, /ready, or equivalent.
Q154: Why do probes need to match the app framework?
Because a probe endpoint must be supported by the application logic.
Q155: How do probes impact rolling updates?
Ready pods receive traffic; unready ones are removed from endpoints while new versions roll out.
Q156: Why do rollout strategies rely on readiness?
Because the Service should route only to stable and ready versions.
Q157: What is a deployment strategy and probe relationship?
The rollout strategy uses readiness to decide when a new pod is eligible for traffic.
Q158: Why is health monitoring from probes useful in debugging?
Because kubelet events and pod conditions reveal why a pod isn’t ready or keeps restarting.
Q159: Why should probes be tracked in resources and events?
Because they indicate bottlenecks or configuration mistakes in the application lifecycle.
Q160: What is a pod lifecycle event log?
It records readiness transitions, deaths, restarts, and liveness failures.
Q161: What is a probe with an exec command using shell?
Example: `["/bin/sh", "-c", "curl -f http://localhost:8080/ready || exit 1"]`
Q162: Which probe type is usually best?
HTTP probes are usually the most readable and maintainable for web services.
Q163: Why is TCP probe not enough for service readiness?
Because it says the port is reachable, not that the application is fully ready to respond.
Q164: Why should startup and readiness differ?
Because startup often includes initialization work while readiness indicates actual service availability.
Q165: What is a readiness signal from an application’s framework?
It could be an actuator health check or a custom endpoint.
Q166: What is a kubelet event when probe fails?
The kubelet may emit warnings or events describing failed readiness or liveness.
Q167: How do operators use probe logs?
They inspect them while debugging restarts or unready pods.
Q168: Why is a too-short probe period a problem?
It can generate excessive checks and load on the app or node.
Q169: Why is a too-long period a problem?
It delays detection of failure or readiness changes.
Q170: What is a good default probe strategy?
Use modest intervals, clear paths, and a meaningful readiness condition aligned to app behavior.
Q171: Why do apps need to control readiness carefully?
Because readiness determines business traffic routing and user impact.
Q172: What is a probe for database migration?
It might wait until migrations are complete before readiness becomes true.
Q173: What is a probe for cache warmup?
It waits until the cache is warmed before allowing traffic.
Q174: What is a probe for certificate loading?
It waits until TLS certificates and keys are loaded before becoming ready.
Q175: What is a probe for network dependencies?
It waits until required external services are reachable before the service accepts traffic.
Q176: Why do startup and readiness often combine?
Because some apps need both time to initialize and a precise signal that the app can handle requests.
Q177: What is a probe returning 500 while startup completes?
That is usually a sign the app is not ready yet or the dependency is still failing.
Q178: Why do readiness and liveness values often differ in real apps?
Because readiness is user-service oriented while liveness is process health oriented.
Q179: What is the relationship between probes and service endpoint updates?
Service endpoints are updated based on readiness. Unready pods are removed from traffic.
Q180: Why does a healthy Service depend on good probes?
Because traffic routing is only as good as the readiness and liveness signals.
Advanced / Expert
Q181: What is a liveness probe after startup?
It becomes the ongoing signal that the process is still alive and functioning.
Q182: Why use startup probes for JVM apps with long warmup?
Because Java apps may take time to start, and liveness should not kill them too early.
Q183: What is a startup probe statement?
A startup probe usually runs until success and then liveness and readiness checks become active.
Q184: What is probe gating in Kubernetes?
Probe gating means the kubelet only checks the readiness and liveness signals at the appropriate lifecycle stage.
Q185: What is a failure to transition from startup to ready?
It usually indicates the app remains broken or is taking too long to initialize.
Q186: Why should readiness not be based only on process presence?
Because a process can be alive while dysfunctional.
Q187: What is an app-level health contract?
It is the contract between app and platform: e.g. "200 means healthy and ready, 503 means not ready".
Q188: Why is a health contract helpful?
It gives operators and developers a clear signal for infrastructure behavior.
Q189: What is a probe and latency trade-off?
Tighter probes detect issues faster but can create noise and false positives; slower probes are less noisy but delay detection.
Q190: What is a failure threshold tuning strategy?
It balances how many failures are tolerated before making the container unhealthy or unready.
Q191: What is a success threshold tuning strategy?
It defines how much evidence is needed before a previously failing probe is considered healthy again.
Q192: Why do readiness and liveness thresholds differ?
Because readiness may recover more quickly while liveness may need a stronger signal before restart.
Q193: What is a probe for dependency outage?
It may check a critical dependency such as a database or message broker.
Q194: Why should dependency checks be selective?
Because checking every downstream dependency may make readiness too sensitive or too expensive.
Q195: What is a readiness endpoint with connection pooling?
It can ensure the app has established necessary connections before being routed.
Q196: Why can probe design be an operational policy?
Because it reflects how the platform expects the application to behave under failure.
Q197: What is a probe bug in production?
It can lead to pods being marked unready or restarted unexpectedly, causing outage or flapping.
Q198: What is flapping?
Flapping is when a pod repeatedly transitions between healthy and unhealthy states.
Q199: Why is flapping dangerous?
It causes unstable traffic routing and complicates investigations.
Q200: What is the main lesson of Kubernetes probes?
Probes are the operational contract between Kubernetes and the application: they define health, readiness, and recovery behavior, and they must reflect real service semantics to avoid outages and premature restarts.