Prometheus & Grafana

Prometheus & Grafana


Beginner

Q1: What is Prometheus?

Prometheus is an open-source metrics monitoring and alerting toolkit.

Q2: What is Grafana?

Grafana is a visualization and dashboard platform for metrics/logs/traces and more.

Q3: Prometheus vs Grafana in one line?

Prometheus stores/queries metrics; Grafana visualizes and explores them.

Q4: What is a metric?

A numeric time-series measurement representing system/application behavior.

Q5: What is a time series?

A sequence of timestamped values for a metric + label set.

Q6: What is label in Prometheus?

Key-value metadata dimension attached to metric series.

Q7: Why labels are useful?

Enable filtering, aggregation, and slicing by dimensions (instance, job, status, etc.).

Q8: What is scrape in Prometheus?

Periodic HTTP pull of metrics endpoint from target.

Q9: What is scrape interval?

Frequency at which Prometheus collects metrics from target.

Q10: What is target in Prometheus?

Endpoint/source being scraped for metrics.

Q11: What is exporter?

Component exposing metrics for systems that don’t natively expose Prometheus format.

Q12: Common exporter example?

Node Exporter for host/system metrics.

Q13: What is /metrics endpoint?

HTTP endpoint exposing metrics in Prometheus text/OpenMetrics format.

Q14: What is Prometheus server?

Core component scraping, storing, and querying metrics.

Q15: What is PromQL?

Prometheus query language for selecting and transforming time series.

Q16: What is instant query?

Query evaluated at a single timestamp.

Q17: What is range query?

Query evaluated across time window for graphing/trends.

Q18: What is alerting rule?

Prometheus expression with condition threshold triggering alert.

Q19: What is Alertmanager?

Component handling alert deduplication, grouping, routing, silencing, notifications.

Q20: Why separate Alertmanager from Prometheus?

Decoupled alert processing and multi-source alert routing.

Q21: What is Grafana data source?

Backend system Grafana queries (Prometheus, Loki, Elasticsearch, etc.).

Q22: What is Grafana dashboard?

Collection of panels visualizing related signals.

Q23: What is Grafana panel?

Single visualization unit (graph, stat, table, heatmap, etc.).

Q24: What is dashboard variable?

Templated parameter for dynamic filtering across panels.

Q25: Why variables matter?

Reusable dashboards across services/environments.

Q26: What is counter metric?

Monotonically increasing metric (resets on restart).

Q27: What is gauge metric?

Value that can go up/down arbitrarily (temperature, queue depth).

Q28: What is histogram metric?

Distribution metric with buckets + count + sum.

Q29: What is summary metric?

Client-side calculated quantiles + count + sum (limitations in aggregation).

Q30: Histogram vs summary quick rule?

Prefer histograms for aggregatable latency percentiles across instances.

Q31: What is cardinality?

Number of unique time series generated by metric label combinations.

Q32: Why high cardinality is dangerous?

Memory/storage/query cost explosion.

Q33: Example high-cardinality anti-pattern?

Labeling metrics with userid/requestid raw values.

Q34: What is retention period?

How long Prometheus stores data locally.

Q35: What is TSDB?

Time series database used internally by Prometheus.

Q36: What is service discovery in Prometheus?

Automatic target discovery from platforms (Kubernetes, Consul, cloud APIs, etc.).

Q37: What is staticconfigs?

Manually defined scrape targets in config.

Q38: What is relabeling?

Transforming/filtering target or metric labels during ingestion pipeline.

Q39: What is recording rule?

Precomputed PromQL result stored as new time series.

Q40: Why use recording rules?

Faster dashboards/alerts and standardized query logic.

Q41: What is scrape timeout?

Max duration allowed for a single scrape request.

Q42: Why timeout matters?

Prevents slow targets from blocking scrape cycle reliability.

Q43: What is up metric?

Built-in metric indicating scrape target availability (1 up, 0 down).

Q44: What is Prometheus pull model advantage?

Simpler target discovery and centralized scraping control.

Q45: Can Prometheus receive pushed metrics?

Usually via Pushgateway for short-lived batch jobs (special case).

Q46: Why Pushgateway should be limited?

Not general event store; misuse can create stale/unowned metrics.

Q47: What is beginner anti-pattern in monitoring?

Only infrastructure metrics, no application/business metrics.

Q48: Another beginner anti-pattern?

Alerting on every metric without SLO context.

Q49: Beginner reliability baseline?

Monitor availability, latency, errors, saturation (golden signals).

Q50: Beginner dashboard baseline?

Service overview: request rate, error rate, latency, resource usage.

Q51: Beginner alert baseline?

Actionable alerts for user-impacting conditions, not noise.

Q52: Why alerts should be actionable?

If no action exists, alert causes fatigue not reliability.

Q53: What is silencing in Alertmanager?

Temporary suppression of matching alerts.

Q54: What is inhibition in Alertmanager?

Suppress lower-priority alerts when higher-priority root alert is firing.

Q55: Why use inhibition?

Reduce alert storms and focus responders.

Q56: What is Grafana Explore?

Ad-hoc query interface for interactive troubleshooting.

Q57: What is annotation in Grafana?

Visual marker on graphs (deploys/incidents/events).

Q58: Why deployment annotations help?

Correlate metric changes with release events quickly.

Q59: Beginner security baseline?

Protect Prometheus/Grafana access and secret data sources.

Q60: Beginner ops baseline?

Back up dashboards/alert rules as code.

Q61: Beginner performance baseline?

Avoid heavy unbounded queries and high-cardinality labels.

Q62: Beginner governance baseline?

Dashboard ownership and alert ownership per service.

Q63: Beginner workflow principle?

Start from user impact, then drill into components.

Q64: Beginner collaboration principle?

Standardize metric naming and label conventions early.

Q65: Beginner best practice?

Prioritize signal quality over dashboard quantity.

Intermediate

Q66: What is rate() in PromQL?

Per-second average increase of counter over range.

Q67: What is irate()?

Instantaneous counter rate based on last two points (spikier).

Q68: rate vs irate usage guideline?

Use rate for alerts/trends; irate for fast-changing visual diagnostics.

Q69: What is increase()?

Total counter increase over time range.

Q70: What is delta()?

Difference between first/last values in range (for gauges).

Q71: What is histogramquantile()?

Estimate quantiles from histogram bucket series.

Q72: Why aggregate before histogramquantile carefully?

Need proper sum by (le,...) semantics to preserve bucket structure.

Q73: What is absent() function?

Detect missing series for expected metrics.

Q74: What is vector matching in PromQL?

Joining series across metrics using label matching rules.

Q75: What are on()/ignoring() modifiers?

Control which labels participate in vector matching.

Q76: What is groupleft/groupright?

Resolve one-to-many/many-to-one vector join cardinality.

Q77: What is recording rule naming convention?

Often colon-separated semantic names (team conventions vary).

Q78: Why standardize recording rules?

Reusable queries and consistent alert semantics.

Q79: What is rule evaluation interval?

How often Prometheus evaluates recording/alert rules.

Q80: What is FOR clause in alert rules?

Condition must remain true for duration before firing.

Q81: Why FOR reduces noise?

Filters transient spikes/flapping conditions.

Q82: What is pending alert state?

Condition met but FOR duration not yet satisfied.

Q83: What is firing state?

Alert active and routed to Alertmanager.

Q84: What is Alertmanager route tree?

Hierarchical routing logic by labels to receivers.

Q85: What is receiver in Alertmanager?

Notification destination/integration (Slack, PagerDuty, email, etc.).

Q86: What is groupby in Alertmanager?

Labels used to batch related alerts into one notification.

Q87: What is groupwait/groupinterval/repeatinterval?

Timing controls for first, subsequent, and repeat notifications.

Q88: What is blackbox exporter?

Probes endpoints externally (HTTP, TCP, DNS, ICMP) for availability/latency.

Q89: When use blackbox probing?

Measure user-like reachability, not only internal app metrics.

Q90: What is kube-state-metrics?

Exports Kubernetes object state metrics (deployments, pods, etc.).

Q91: Node exporter vs kube-state-metrics?

Node exporter: host OS metrics; kube-state-metrics: Kubernetes object metadata/state.

Q92: What is ServiceMonitor (Prometheus Operator)?

CRD defining scrape config for Kubernetes services.

Q93: What is PodMonitor?

CRD defining scrape config for pods directly.

Q94: Why Prometheus Operator?

Simplifies Kubernetes-native management of Prometheus stack.

Q95: What is remotewrite?

Forward metrics to remote long-term storage backend.

Q96: Why use remotewrite?

Long retention, global query, durable storage beyond local Prometheus.

Q97: What is remoteread?

Query external storage through Prometheus interface (backend dependent).

Q98: What is federation in Prometheus?

Prometheus scraping aggregated metrics from other Prometheus servers.

Q99: Federation vs remotewrite?

Federation pulls selected series; remotewrite streams samples out.

Q100: What is downsampling concept?

Store lower-resolution historical data for long-term cost efficiency.

Q101: What is exemplars concept?

Attach trace/sample IDs to metrics points for metric-to-trace correlation.

Q102: Why exemplars useful?

Faster root-cause pivot from latency spikes to traces.

Q103: What is RED method?

Monitor Rate, Errors, Duration for request-driven services.

Q104: What is USE method?

Monitor Utilization, Saturation, Errors for resources.

Q105: What is SLI?

Service Level Indicator: measured reliability metric.

Q106: What is SLO?

Target objective for SLI over time.

Q107: What is error budget?

Allowed unreliability derived from SLO target.

Q108: Why alert on error budget burn?

Align alerts with user impact and reliability goals.

Q109: What is multi-window multi-burn-rate alerting?

Combines short/long windows for fast and stable SLO breach detection.

Q110: What is intermediate anti-pattern?

Alerting directly on CPU/memory without service-impact context.

Q111: Better alert strategy?

Symptoms (user impact) first, causes second.

Q112: What is dashboard anti-pattern?

Dense dashboard with no narrative hierarchy.

Q113: Better dashboard structure?

Overview -> drilldown -> component diagnostics.

Q114: What is Grafana folder RBAC?

Access control grouping dashboards by team/domain sensitivity.

Q115: What is provisioning in Grafana?

Managing dashboards/datasources/alerts via files/APIs as code.

Q116: Why dashboard-as-code?

Version control, review, reproducibility, rollback.

Q117: What is intermediate security baseline?

AuthN/AuthZ, datasource credential scoping, audit logging.

Q118: What is intermediate reliability baseline?

HA alert routing and tested notification paths.

Q119: What is intermediate performance baseline?

Recording rules and constrained query ranges.

Q120: Intermediate maturity signal?

Team can trace alert -> dashboard -> runbook quickly.

Q121: What is runbook link in alert?

Operational guide URL embedded in alert annotations.

Q122: Why runbooks reduce MTTR?

Immediate context and remediation steps for responders.

Q123: What is metric relabel drop use?

Discard unnecessary/high-cardinality series at scrape time.

Q124: What is scrape job sharding concept?

Split target load across Prometheus instances.

Q125: What is intermediate ops principle?

Treat observability configs as production code with CI checks.

Q126: What is intermediate governance principle?

Define naming/labeling/alert standards org-wide.

Q127: What is intermediate cost principle?

Continuously prune low-value metrics and dashboards.

Q128: What is intermediate architecture principle?

Separate collection, storage, alerting, and visualization responsibilities.

Q129: What is intermediate collaboration principle?

SRE + dev teams co-own service-level observability.

Q130: Intermediate best practice?

Optimize for actionable clarity, not metric volume.

Advanced

Q131: What is Prometheus HA pair pattern?

Two independent Prometheus servers scrape same targets for redundancy.

Q132: Why HA Prometheus still needs care?

Duplicate alerts/samples handling requires Alertmanager/remote backend dedup logic.

Q133: What is global metrics architecture challenge?

Balancing local autonomy with centralized visibility and cost.

Q134: What is Thanos/Cortex/Mimir-style role?

Horizontally scalable long-term Prometheus-compatible storage/query layers.

Q135: Why add long-term backend?

Years of retention, global querying, and durable storage.

Q136: What is query frontend role (in scalable backends)?

Caching, splitting, and optimizing query execution.

Q137: What is compaction in TSDB systems?

Merge blocks/chunks for storage efficiency and query performance.

Q138: What is cardinality explosion incident?

Rapid uncontrolled series growth causing memory/query failures.

Q139: Cardinality mitigation playbook?

Identify top offenders, relabel/drop, redesign labels, enforce budgets.

Q140: What is cardinality budget?

Per-team/service cap on allowed active series dimensions.

Q141: Why budget cardinality explicitly?

Prevents shared observability platform exhaustion.

Q142: What is staleness in Prometheus?

Series marked stale when targets disappear/sample stream stops.

Q143: Why staleness semantics matter?

Avoid misleading joins/aggregations during target churn.

Q144: What is native histogram concept?

Emerging histogram representation improving accuracy/performance tradeoffs (version/ecosystem dependent).

Q145: What is remotewrite backpressure risk?

Queue buildup/data loss risk during remote endpoint slowdown.

Q146: Mitigation for remotewrite outages?

Queue tuning, WAL durability awareness, backend HA, alerting.

Q147: What is WAL in Prometheus?

Write-ahead log for durability and crash recovery.

Q148: What is scrape jitter strategy?

Stagger scrape timings to reduce synchronized load spikes.

Q149: What is tenant isolation in observability?

Separate data access/quotas/routing for teams/customers.

Q150: What is noisy neighbor problem in metrics backend?

One tenant’s high-cardinality/high-query load degrades others.

Q151: Mitigation for noisy neighbors?

Per-tenant quotas, query limits, isolated ingestion paths.

Q152: What is query cost governance?

Limit expensive PromQL patterns and enforce timeout/max samples.

Q153: What is recording-rule anti-pattern at scale?

Precomputing too many low-value series increasing storage load.

Q154: Better recording rule strategy?

Precompute only expensive/high-traffic/high-value queries.

Q155: What is alert flood scenario?

Many related alerts fire simultaneously during major outage.

Q156: Alert flood mitigation?

Inhibition, grouping, symptom-based paging, cause alerts as tickets.

Q157: What is paging philosophy for SRE?

Page on user-impacting symptoms, not every internal anomaly.

Q158: What is burn-rate alert tuning challenge?

Too sensitive causes noise; too lax delays detection.

Q159: How calibrate burn-rate alerts?

Historical incident replay + game days + SLO error budget policy.

Q160: What is observability security risk?

Metrics may leak sensitive labels/data if instrumentation is careless.

Q161: How prevent sensitive data leakage in metrics?

Never include PII/secrets in labels/metric values.

Q162: What is mTLS/auth for scrape endpoints?

Secure transport/authentication between Prometheus and targets.

Q163: What is Grafana enterprise-scale challenge?

Dashboard sprawl, inconsistent ownership, variable query quality.

Q164: Grafana sprawl mitigation?

Folder taxonomy, review process, dashboard lifecycle policies.

Q165: What is dashboard lifecycle management?

Create, review, deprecate, archive dashboards by usage/ownership.

Q166: What is usage analytics in Grafana?

Track dashboard/query usage to prune low-value assets.

Q167: What is incident timeline correlation pattern?

Overlay deploys, config changes, and alerts on metrics timeline.

Q168: Why correlate change events?

Speeds root cause identification and rollback decisions.

Q169: What is synthetic + real-user monitoring relation?

Synthetic probes baseline availability; RUM captures real client experience.

Q170: Why combine both?

Broader coverage of controlled and real-world behavior.

Q171: What is chaos engineering role in observability?

Validate alerts/dashboards detect expected failure modes.

Q172: What is observability game day?

Practice incidents to test signal quality and runbooks.

Q173: What is compliance evidence from monitoring stack?

Audit logs, alert history, SLO reports, access records.

Q174: What is DR plan for observability platform?

Backup configs, replicate storage, restore dashboards/alerts quickly.

Q175: What is final reliability principle?

Monitoring must remain available during incidents, not fail first.

Q176: What is final signal-quality principle?

Fewer actionable alerts beat many noisy alerts.

Q177: What is final cost principle?

Every metric/label/query should justify its operational value.

Q178: What is final security principle?

Protect telemetry pipelines and prevent sensitive data exposure.

Q179: What is final governance principle?

Standardize instrumentation, labels, and alert semantics organization-wide.

Q180: What is final architecture principle?

Design observability as layered system: collect, store, alert, visualize.

Q181: What is final operations principle?

Continuously test alerts, runbooks, and escalation paths.

Q182: What is final collaboration principle?

Developers and SREs share ownership of service observability.

Q183: What is final scalability principle?

Plan cardinality and query growth before incidents force redesign.

Q184: What is final product principle?

Dashboards are operational products with users and maintenance lifecycle.

Q185: Final maturity principle?

Prometheus/Grafana excellence is actionable, scalable, and SLO-driven observability.

Bonus: Minimal Prometheus Alert Rule Example

groups:
  - name: service-alerts
    rules:
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{job="api",status=~"5.."}[5m]))
          /
          sum(rate(http_requests_total{job="api"}[5m])) > 0.05
        for: 10m
        labels:
          severity: page
          service: api
        annotations:
          summary: "API error rate is above 5%"
          description: "5xx ratio exceeded 5% for 10 minutes."