Prometheus & Grafana
Prometheus & Grafana
Beginner
Q1: What is Prometheus?
Prometheus is an open-source metrics monitoring and alerting toolkit.
Q2: What is Grafana?
Grafana is a visualization and dashboard platform for metrics/logs/traces and more.
Q3: Prometheus vs Grafana in one line?
Prometheus stores/queries metrics; Grafana visualizes and explores them.
Q4: What is a metric?
A numeric time-series measurement representing system/application behavior.
Q5: What is a time series?
A sequence of timestamped values for a metric + label set.
Q6: What is label in Prometheus?
Key-value metadata dimension attached to metric series.
Q7: Why labels are useful?
Enable filtering, aggregation, and slicing by dimensions (instance, job, status, etc.).
Q8: What is scrape in Prometheus?
Periodic HTTP pull of metrics endpoint from target.
Q9: What is scrape interval?
Frequency at which Prometheus collects metrics from target.
Q10: What is target in Prometheus?
Endpoint/source being scraped for metrics.
Q11: What is exporter?
Component exposing metrics for systems that don’t natively expose Prometheus format.
Q12: Common exporter example?
Node Exporter for host/system metrics.
Q13: What is /metrics endpoint?
HTTP endpoint exposing metrics in Prometheus text/OpenMetrics format.
Q14: What is Prometheus server?
Core component scraping, storing, and querying metrics.
Q15: What is PromQL?
Prometheus query language for selecting and transforming time series.
Q16: What is instant query?
Query evaluated at a single timestamp.
Q17: What is range query?
Query evaluated across time window for graphing/trends.
Q18: What is alerting rule?
Prometheus expression with condition threshold triggering alert.
Q19: What is Alertmanager?
Component handling alert deduplication, grouping, routing, silencing, notifications.
Q20: Why separate Alertmanager from Prometheus?
Decoupled alert processing and multi-source alert routing.
Q21: What is Grafana data source?
Backend system Grafana queries (Prometheus, Loki, Elasticsearch, etc.).
Q22: What is Grafana dashboard?
Collection of panels visualizing related signals.
Q23: What is Grafana panel?
Single visualization unit (graph, stat, table, heatmap, etc.).
Q24: What is dashboard variable?
Templated parameter for dynamic filtering across panels.
Q25: Why variables matter?
Reusable dashboards across services/environments.
Q26: What is counter metric?
Monotonically increasing metric (resets on restart).
Q27: What is gauge metric?
Value that can go up/down arbitrarily (temperature, queue depth).
Q28: What is histogram metric?
Distribution metric with buckets + count + sum.
Q29: What is summary metric?
Client-side calculated quantiles + count + sum (limitations in aggregation).
Q30: Histogram vs summary quick rule?
Prefer histograms for aggregatable latency percentiles across instances.
Q31: What is cardinality?
Number of unique time series generated by metric label combinations.
Q32: Why high cardinality is dangerous?
Memory/storage/query cost explosion.
Q33: Example high-cardinality anti-pattern?
Labeling metrics with userid/requestid raw values.
Q34: What is retention period?
How long Prometheus stores data locally.
Q35: What is TSDB?
Time series database used internally by Prometheus.
Q36: What is service discovery in Prometheus?
Automatic target discovery from platforms (Kubernetes, Consul, cloud APIs, etc.).
Q37: What is staticconfigs?
Manually defined scrape targets in config.
Q38: What is relabeling?
Transforming/filtering target or metric labels during ingestion pipeline.
Q39: What is recording rule?
Precomputed PromQL result stored as new time series.
Q40: Why use recording rules?
Faster dashboards/alerts and standardized query logic.
Q41: What is scrape timeout?
Max duration allowed for a single scrape request.
Q42: Why timeout matters?
Prevents slow targets from blocking scrape cycle reliability.
Q43: What is up metric?
Built-in metric indicating scrape target availability (1 up, 0 down).
Q44: What is Prometheus pull model advantage?
Simpler target discovery and centralized scraping control.
Q45: Can Prometheus receive pushed metrics?
Usually via Pushgateway for short-lived batch jobs (special case).
Q46: Why Pushgateway should be limited?
Not general event store; misuse can create stale/unowned metrics.
Q47: What is beginner anti-pattern in monitoring?
Only infrastructure metrics, no application/business metrics.
Q48: Another beginner anti-pattern?
Alerting on every metric without SLO context.
Q49: Beginner reliability baseline?
Monitor availability, latency, errors, saturation (golden signals).
Q50: Beginner dashboard baseline?
Service overview: request rate, error rate, latency, resource usage.
Q51: Beginner alert baseline?
Actionable alerts for user-impacting conditions, not noise.
Q52: Why alerts should be actionable?
If no action exists, alert causes fatigue not reliability.
Q53: What is silencing in Alertmanager?
Temporary suppression of matching alerts.
Q54: What is inhibition in Alertmanager?
Suppress lower-priority alerts when higher-priority root alert is firing.
Q55: Why use inhibition?
Reduce alert storms and focus responders.
Q56: What is Grafana Explore?
Ad-hoc query interface for interactive troubleshooting.
Q57: What is annotation in Grafana?
Visual marker on graphs (deploys/incidents/events).
Q58: Why deployment annotations help?
Correlate metric changes with release events quickly.
Q59: Beginner security baseline?
Protect Prometheus/Grafana access and secret data sources.
Q60: Beginner ops baseline?
Back up dashboards/alert rules as code.
Q61: Beginner performance baseline?
Avoid heavy unbounded queries and high-cardinality labels.
Q62: Beginner governance baseline?
Dashboard ownership and alert ownership per service.
Q63: Beginner workflow principle?
Start from user impact, then drill into components.
Q64: Beginner collaboration principle?
Standardize metric naming and label conventions early.
Q65: Beginner best practice?
Prioritize signal quality over dashboard quantity.
Intermediate
Q66: What is rate() in PromQL?
Per-second average increase of counter over range.
Q67: What is irate()?
Instantaneous counter rate based on last two points (spikier).
Q68: rate vs irate usage guideline?
Use rate for alerts/trends; irate for fast-changing visual diagnostics.
Q69: What is increase()?
Total counter increase over time range.
Q70: What is delta()?
Difference between first/last values in range (for gauges).
Q71: What is histogramquantile()?
Estimate quantiles from histogram bucket series.
Q72: Why aggregate before histogramquantile carefully?
Need proper sum by (le,...) semantics to preserve bucket structure.
Q73: What is absent() function?
Detect missing series for expected metrics.
Q74: What is vector matching in PromQL?
Joining series across metrics using label matching rules.
Q75: What are on()/ignoring() modifiers?
Control which labels participate in vector matching.
Q76: What is groupleft/groupright?
Resolve one-to-many/many-to-one vector join cardinality.
Q77: What is recording rule naming convention?
Often colon-separated semantic names (team conventions vary).
Q78: Why standardize recording rules?
Reusable queries and consistent alert semantics.
Q79: What is rule evaluation interval?
How often Prometheus evaluates recording/alert rules.
Q80: What is FOR clause in alert rules?
Condition must remain true for duration before firing.
Q81: Why FOR reduces noise?
Filters transient spikes/flapping conditions.
Q82: What is pending alert state?
Condition met but FOR duration not yet satisfied.
Q83: What is firing state?
Alert active and routed to Alertmanager.
Q84: What is Alertmanager route tree?
Hierarchical routing logic by labels to receivers.
Q85: What is receiver in Alertmanager?
Notification destination/integration (Slack, PagerDuty, email, etc.).
Q86: What is groupby in Alertmanager?
Labels used to batch related alerts into one notification.
Q87: What is groupwait/groupinterval/repeatinterval?
Timing controls for first, subsequent, and repeat notifications.
Q88: What is blackbox exporter?
Probes endpoints externally (HTTP, TCP, DNS, ICMP) for availability/latency.
Q89: When use blackbox probing?
Measure user-like reachability, not only internal app metrics.
Q90: What is kube-state-metrics?
Exports Kubernetes object state metrics (deployments, pods, etc.).
Q91: Node exporter vs kube-state-metrics?
Node exporter: host OS metrics; kube-state-metrics: Kubernetes object metadata/state.
Q92: What is ServiceMonitor (Prometheus Operator)?
CRD defining scrape config for Kubernetes services.
Q93: What is PodMonitor?
CRD defining scrape config for pods directly.
Q94: Why Prometheus Operator?
Simplifies Kubernetes-native management of Prometheus stack.
Q95: What is remotewrite?
Forward metrics to remote long-term storage backend.
Q96: Why use remotewrite?
Long retention, global query, durable storage beyond local Prometheus.
Q97: What is remoteread?
Query external storage through Prometheus interface (backend dependent).
Q98: What is federation in Prometheus?
Prometheus scraping aggregated metrics from other Prometheus servers.
Q99: Federation vs remotewrite?
Federation pulls selected series; remotewrite streams samples out.
Q100: What is downsampling concept?
Store lower-resolution historical data for long-term cost efficiency.
Q101: What is exemplars concept?
Attach trace/sample IDs to metrics points for metric-to-trace correlation.
Q102: Why exemplars useful?
Faster root-cause pivot from latency spikes to traces.
Q103: What is RED method?
Monitor Rate, Errors, Duration for request-driven services.
Q104: What is USE method?
Monitor Utilization, Saturation, Errors for resources.
Q105: What is SLI?
Service Level Indicator: measured reliability metric.
Q106: What is SLO?
Target objective for SLI over time.
Q107: What is error budget?
Allowed unreliability derived from SLO target.
Q108: Why alert on error budget burn?
Align alerts with user impact and reliability goals.
Q109: What is multi-window multi-burn-rate alerting?
Combines short/long windows for fast and stable SLO breach detection.
Q110: What is intermediate anti-pattern?
Alerting directly on CPU/memory without service-impact context.
Q111: Better alert strategy?
Symptoms (user impact) first, causes second.
Q112: What is dashboard anti-pattern?
Dense dashboard with no narrative hierarchy.
Q113: Better dashboard structure?
Overview -> drilldown -> component diagnostics.
Q114: What is Grafana folder RBAC?
Access control grouping dashboards by team/domain sensitivity.
Q115: What is provisioning in Grafana?
Managing dashboards/datasources/alerts via files/APIs as code.
Q116: Why dashboard-as-code?
Version control, review, reproducibility, rollback.
Q117: What is intermediate security baseline?
AuthN/AuthZ, datasource credential scoping, audit logging.
Q118: What is intermediate reliability baseline?
HA alert routing and tested notification paths.
Q119: What is intermediate performance baseline?
Recording rules and constrained query ranges.
Q120: Intermediate maturity signal?
Team can trace alert -> dashboard -> runbook quickly.
Q121: What is runbook link in alert?
Operational guide URL embedded in alert annotations.
Q122: Why runbooks reduce MTTR?
Immediate context and remediation steps for responders.
Q123: What is metric relabel drop use?
Discard unnecessary/high-cardinality series at scrape time.
Q124: What is scrape job sharding concept?
Split target load across Prometheus instances.
Q125: What is intermediate ops principle?
Treat observability configs as production code with CI checks.
Q126: What is intermediate governance principle?
Define naming/labeling/alert standards org-wide.
Q127: What is intermediate cost principle?
Continuously prune low-value metrics and dashboards.
Q128: What is intermediate architecture principle?
Separate collection, storage, alerting, and visualization responsibilities.
Q129: What is intermediate collaboration principle?
SRE + dev teams co-own service-level observability.
Q130: Intermediate best practice?
Optimize for actionable clarity, not metric volume.
Advanced
Q131: What is Prometheus HA pair pattern?
Two independent Prometheus servers scrape same targets for redundancy.
Q132: Why HA Prometheus still needs care?
Duplicate alerts/samples handling requires Alertmanager/remote backend dedup logic.
Q133: What is global metrics architecture challenge?
Balancing local autonomy with centralized visibility and cost.
Q134: What is Thanos/Cortex/Mimir-style role?
Horizontally scalable long-term Prometheus-compatible storage/query layers.
Q135: Why add long-term backend?
Years of retention, global querying, and durable storage.
Q136: What is query frontend role (in scalable backends)?
Caching, splitting, and optimizing query execution.
Q137: What is compaction in TSDB systems?
Merge blocks/chunks for storage efficiency and query performance.
Q138: What is cardinality explosion incident?
Rapid uncontrolled series growth causing memory/query failures.
Q139: Cardinality mitigation playbook?
Identify top offenders, relabel/drop, redesign labels, enforce budgets.
Q140: What is cardinality budget?
Per-team/service cap on allowed active series dimensions.
Q141: Why budget cardinality explicitly?
Prevents shared observability platform exhaustion.
Q142: What is staleness in Prometheus?
Series marked stale when targets disappear/sample stream stops.
Q143: Why staleness semantics matter?
Avoid misleading joins/aggregations during target churn.
Q144: What is native histogram concept?
Emerging histogram representation improving accuracy/performance tradeoffs (version/ecosystem dependent).
Q145: What is remotewrite backpressure risk?
Queue buildup/data loss risk during remote endpoint slowdown.
Q146: Mitigation for remotewrite outages?
Queue tuning, WAL durability awareness, backend HA, alerting.
Q147: What is WAL in Prometheus?
Write-ahead log for durability and crash recovery.
Q148: What is scrape jitter strategy?
Stagger scrape timings to reduce synchronized load spikes.
Q149: What is tenant isolation in observability?
Separate data access/quotas/routing for teams/customers.
Q150: What is noisy neighbor problem in metrics backend?
One tenant’s high-cardinality/high-query load degrades others.
Q151: Mitigation for noisy neighbors?
Per-tenant quotas, query limits, isolated ingestion paths.
Q152: What is query cost governance?
Limit expensive PromQL patterns and enforce timeout/max samples.
Q153: What is recording-rule anti-pattern at scale?
Precomputing too many low-value series increasing storage load.
Q154: Better recording rule strategy?
Precompute only expensive/high-traffic/high-value queries.
Q155: What is alert flood scenario?
Many related alerts fire simultaneously during major outage.
Q156: Alert flood mitigation?
Inhibition, grouping, symptom-based paging, cause alerts as tickets.
Q157: What is paging philosophy for SRE?
Page on user-impacting symptoms, not every internal anomaly.
Q158: What is burn-rate alert tuning challenge?
Too sensitive causes noise; too lax delays detection.
Q159: How calibrate burn-rate alerts?
Historical incident replay + game days + SLO error budget policy.
Q160: What is observability security risk?
Metrics may leak sensitive labels/data if instrumentation is careless.
Q161: How prevent sensitive data leakage in metrics?
Never include PII/secrets in labels/metric values.
Q162: What is mTLS/auth for scrape endpoints?
Secure transport/authentication between Prometheus and targets.
Q163: What is Grafana enterprise-scale challenge?
Dashboard sprawl, inconsistent ownership, variable query quality.
Q164: Grafana sprawl mitigation?
Folder taxonomy, review process, dashboard lifecycle policies.
Q165: What is dashboard lifecycle management?
Create, review, deprecate, archive dashboards by usage/ownership.
Q166: What is usage analytics in Grafana?
Track dashboard/query usage to prune low-value assets.
Q167: What is incident timeline correlation pattern?
Overlay deploys, config changes, and alerts on metrics timeline.
Q168: Why correlate change events?
Speeds root cause identification and rollback decisions.
Q169: What is synthetic + real-user monitoring relation?
Synthetic probes baseline availability; RUM captures real client experience.
Q170: Why combine both?
Broader coverage of controlled and real-world behavior.
Q171: What is chaos engineering role in observability?
Validate alerts/dashboards detect expected failure modes.
Q172: What is observability game day?
Practice incidents to test signal quality and runbooks.
Q173: What is compliance evidence from monitoring stack?
Audit logs, alert history, SLO reports, access records.
Q174: What is DR plan for observability platform?
Backup configs, replicate storage, restore dashboards/alerts quickly.
Q175: What is final reliability principle?
Monitoring must remain available during incidents, not fail first.
Q176: What is final signal-quality principle?
Fewer actionable alerts beat many noisy alerts.
Q177: What is final cost principle?
Every metric/label/query should justify its operational value.
Q178: What is final security principle?
Protect telemetry pipelines and prevent sensitive data exposure.
Q179: What is final governance principle?
Standardize instrumentation, labels, and alert semantics organization-wide.
Q180: What is final architecture principle?
Design observability as layered system: collect, store, alert, visualize.
Q181: What is final operations principle?
Continuously test alerts, runbooks, and escalation paths.
Q182: What is final collaboration principle?
Developers and SREs share ownership of service observability.
Q183: What is final scalability principle?
Plan cardinality and query growth before incidents force redesign.
Q184: What is final product principle?
Dashboards are operational products with users and maintenance lifecycle.
Q185: Final maturity principle?
Prometheus/Grafana excellence is actionable, scalable, and SLO-driven observability.
Bonus: Minimal Prometheus Alert Rule Example
groups:
- name: service-alerts
rules:
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{job="api",status=~"5.."}[5m]))
/
sum(rate(http_requests_total{job="api"}[5m])) > 0.05
for: 10m
labels:
severity: page
service: api
annotations:
summary: "API error rate is above 5%"
description: "5xx ratio exceeded 5% for 10 minutes."