Elasticsearch & Logstash & Kibana
Elasticsearch & Logstash & Kibana
Beginner
Q1: What is ELK?
ELK is a popular stack consisting of Elasticsearch, Logstash, and Kibana used for log collection, indexing, searching, and visualization.
Q2: What is Elasticsearch?
Elasticsearch is a distributed search and analytics engine that stores data in JSON documents and supports full-text and structured queries.
Q3: What is Logstash?
Logstash is a data processing pipeline that collects, transforms, and forwards logs or events to Elasticsearch.
Q4: What is Kibana?
Kibana is the visualization and exploration layer for Elasticsearch data.
Q5: What is an ELK stack used for?
It is commonly used for centralized logging, operational analytics, metrics exploration, and troubleshooting.
Q6: Why use ELK?
Because it provides a scalable, queryable way to store and inspect large amounts of logs and events.
Q7: What is an Elasticsearch index?
An index is a logical collection of documents in Elasticsearch.
Q8: What is a document?
A document is a JSON object stored in Elasticsearch.
Q9: What is a field?
A field is a key in a JSON document, such as timestamp, status, or message.
Q10: What is a shard?
A shard is a partition of an index used to distribute data across nodes.
Q11: What is a replica?
A replica is a copy of a shard used for redundancy and availability.
Q12: Why are replicas important?
They improve fault tolerance and allow failover if a node goes down.
Q13: What is a cluster in Elasticsearch?
A cluster is a set of Elasticsearch nodes that work together to store and index data.
Q14: What is a node in Elasticsearch?
A node is a single Elasticsearch instance running in a cluster.
Q15: What is a master node?
A master node is responsible for cluster state management and orchestration.
Q16: What is a data node?
A data node stores and indexes the actual data.
Q17: What is ingestion?
Ingestion is the process of sending data into Elasticsearch.
Q18: What is indexing?
Indexing is the process of parsing, storing, and making documents searchable.
Q19: What is searching?
Searching means querying Elasticsearch for documents or aggregations.
Q20: What is an analyzer?
An analyzer is a component that tokenizes and normalizes text before indexing or searching.
Q21: What is full-text search?
Full-text search indexes text and enables keyword matching, relevance scoring, and tokenization.
Q22: What is a mapping?
A mapping defines the fields and data types in an index.
Q23: What is a type in Elasticsearch?
Earlier versions used document types, but newer versions use indices and mappings without types.
Q24: Why is mapping important?
Because it defines how fields are indexed and searched.
Q25: What is a string field?
A string field usually uses text or keyword mappings.
Q26: What is a keyword field?
A keyword field is used for exact-match filtering and aggregations.
Q27: What is a text field?
A text field is analyzed for full-text search.
Q28: What is a numeric field?
A numeric field stores numbers such as integer, float, or long.
Q29: What is a date field?
A date field stores timestamps and supports date range queries.
Q30: What is a boolean field?
A boolean field stores true/false values.
Q31: What is a nested field?
A nested field stores arrays of objects with independent query semantics.
Q32: What is a geo field?
A geo field stores geographic coordinates or shapes.
Q33: What is a query?
A query is a request for documents that match certain conditions.
Q34: What is a filter?
A filter is usually a query that can be cached and used for exact matching.
Q35: What is a bool query?
A bool query combines clauses like must, should, mustnot, and filter.
Q36: What is a match query?
A match query performs full-text matching on a field.
Q37: What is a term query?
A term query matches exact values, often used for keyword fields.
Q38: What is a range query?
A range query matches values between lower and upper bounds, often used for timestamps and numeric fields.
Q39: What is a wildcard query?
A wildcard query matches text patterns using wildcards like * or ?.
Q40: What is a regex query?
A regex query matches data with regular-expression patterns.
Q41: What is a sort?
Sorting controls the order of search results.
Q42: What is a score?
A score is the relevance score assigned to a document by Elasticsearch.
Q43: What is relevance?
Relevance is how well a document matches a query relative to others.
Q44: What is a composite aggregation?
A composite aggregation combines bucket results from multiple dimensions.
Q45: What is a bucket aggregation?
A bucket aggregation groups documents by certain criteria.
Q46: What is a metric aggregation?
A metric aggregation computes statistics like average, sum, min, max, or percentiles.
Q47: What is a histogram?
A histogram groups documents by a time or value interval.
Q48: What is a date histogram?
A date histogram groups data into time buckets such as minute, hour, or day.
Q49: What is a term aggregation?
A term aggregation groups documents by a field value.
Q50: What is a stats aggregation?
A stats aggregation calculates count, min, max, sum, and average for a field.
Q51: What is a percentile aggregation?
A percentile aggregation shows thresholds like p95 or p99 of a metric.
Q52: What is a log event?
A log event is a single record of an event or log message.
Q53: What is structured logging?
Structured logging stores logs as structured JSON rather than plain text.
Q54: Why is structured logging useful?
It makes log filtering and analysis easier and more precise.
Q55: What is Logstash pipeline?
A Logstash pipeline is a sequence of inputs, filters, and outputs.
Q56: What is an input plugin?
An input plugin reads data from a source, such as file, syslog, beats, Kafka, or HTTP.
Q57: What is a filter plugin?
A filter plugin transforms or enriches the event.
Q58: What is an output plugin?
An output plugin sends data to a destination such as Elasticsearch.
Q59: What is a Beats agent?
Beats is a lightweight data shipper used to collect logs and metrics from hosts or applications.
Q60: What is Filebeat?
Filebeat ships log files to Logstash or Elasticsearch.
Q61: What is Metricbeat?
Metricbeat ships system and application metrics.
Q62: What is Heartbeat?
Heartbeat checks service availability and liveness.
Q63: What is Packetbeat?
Packetbeat captures network packet data.
Q64: Why use Beats with ELK?
Because Beats are lightweight, agent-based log shippers.
Q65: What is an Elasticsearch index pattern?
An index pattern tells Kibana which indices to search and visualize.
Q66: What is Kibana Discover?
Kibana Discover lets users search and inspect log data interactively.
Q67: What is Kibana Dashboard?
A Dashboard is a set of visualizations showing metrics, logs, and trends.
Q68: What is Kibana Visualization?
A Visualization is a chart or graph built from Elasticsearch queries.
Q69: What is Kibana Search?
Search in Kibana is based on Elasticsearch queries and filters.
Q70: What is a logstash input?
It is how Logstash collects log streams, such as file input or Beats input.
Q71: What is a logstash codec?
A codec defines how data is encoded or decoded in the pipeline.
Q72: What is a grok filter?
A grok filter parses unstructured text using patterns to extract fields.
Q73: Why is grok important?
Because many logs are human-readable text and need structured extraction.
Q74: What is a mutate filter?
A mutate filter changes fields, adds, renames, or removes values.
Q75: What is a date filter?
A date filter parses dates into proper timestamp fields.
Q76: What is a geoip filter?
A geoip filter enriches events with geographic metadata based on IP addresses.
Q77: What is a drop filter?
A drop filter can drop certain events based on conditions.
Q78: What is a condition in Logstash?
A condition often uses if or else blocks based on field values.
Q79: What is an output to Elasticsearch?
The output plugin sends processed log data to Elasticsearch for indexing.
Q80: Why is ELK often paired with Docker or Kubernetes?
Because containers and clusters create huge volumes of logs and metrics that need centralized processing.
Q81: What are common log sources?
Application logs, system logs, access logs, error logs, and security logs.
Q82: Why centralize logs?
Because distributed systems produce logs across many hosts and containers.
Q83: What is log volume?
Log volume is the amount of log data generated by services and infrastructure.
Q84: What is log retention?
Retention is how long logs are kept before being deleted or archived.
Q85: Why do logs need retention policies?
Because keeping all logs indefinitely is expensive and may violate compliance needs.
Q86: What is a log pipeline?
A log pipeline is the path from source to storage and visualization.
Q87: What is indexing performance?
It is the speed at which Elasticsearch can ingest and make data searchable.
Q88: What is cluster health?
Cluster health indicates whether the Elasticsearch cluster is green, yellow, or red.
Q89: What is a green cluster?
All primary shards are assigned and available.
Q90: What is a yellow cluster?
All primary shards are assigned but some replicas are unassigned.
Q91: What is a red cluster?
At least one primary shard is unassigned, meaning some data is unavailable.
Q92: Why is cluster health important?
It indicates whether the cluster is healthy and data is being indexed safely.
Q93: What is a shard allocation?
It is the assignment of shards to nodes in the cluster.
Q94: Why is shard balancing important?
To spread load across nodes and maximize availability.
Q95: What is a hot/warm/cold architecture?
It describes tiered storage strategies for different service levels or retention needs.
Q96: What is an index template?
An index template defines default mappings and settings for new indices.
Q97: Why use index templates?
They reduce mistakes and keep indices consistent.
Q98: What is a rollover index?
A rollover index is a new index created when an existing one reaches size or time thresholds.
Q99: Why rollover indices?
To keep indices manageable and improve search performance.
Q100: What is a time-based index pattern?
A name pattern such as logs-2026.10.* organizes documents by time.
Intermediate
Q101: What is Elasticsearch bulk indexing?
Bulk indexing sends many documents in one request for better performance.
Q102: Why do bulk requests matter?
They improve ingestion throughput and reduce overhead.
Q103: What is refresh interval?
The refresh interval controls when new documents become visible to searches.
Q104: Why is refresh interval important?
Because it trades ingestion speed for search latency.
Q105: What is near real-time search?
Elasticsearch search is near real-time, meaning documents become visible shortly after indexing.
Q106: What is document ingestion throughput?
It is the volume of documents per unit time that Elasticsearch can ingest.
Q107: What is a Lucene index?
Lucene is the search engine beneath Elasticsearch.
Q108: Why does Elasticsearch rely on Lucene?
Because Lucene provides the underlying indexing, search, and scoring capabilities.
Q109: What is a query DSL?
The query DSL is the Elasticsearch domain-specific language for writing queries.
Q110: What is a filter context?
A filter context is used for exact matching and often caches results.
Q111: What is a query context?
A query context computes relevance and scoring.
Q112: Why separate filter and query contexts?
Because filters are efficient and often reusable, while queries compute relevance.
Q113: What is a script query?
A script query runs a script to decide whether a document matches.
Q114: Why use scripts sparingly?
They are slower and more complex than standard queries.
Q115: What is a function score query?
A function score query modifies the relevance score using custom functions.
Q116: What is a geospatial query?
A geospatial query filters or scores documents by geographic location.
Q117: What is a nested query?
A nested query searches within nested objects.
Q118: Why is nested indexing useful?
Because it supports document structures where arrays of objects are semantically grouped.
Q119: What is an index alias?
An alias points to one or more indices, which helps with time-based index rollover and zero-downtime updates.
Q120: Why do aliases help in production?
They let you redirect searches to new indices without changing queries.
Q121: What is a mapping conflict?
A mapping conflict occurs when a field is mapped differently across indices or updates.
Q122: Why do mapping conflicts matter?
They can break indexing or cause unexpected search behavior.
Q123: What is `text` vs `keyword` mapping?
Text is meant for full-text search; keyword is for exact matching and sorting.
Q124: Why is analysis important for text search?
Because tokens are normalized and lowercased so searches are more useful.
Q125: What is stemming?
Stemming reduces words to their root form, such as "running" to "run".
Q126: What is tokenization?
Tokenization splits text into searchable pieces or tokens.
Q127: What is a stop word?
A stop word is a common word removed from indexing, such as "the" or "and".
Q128: Why is field data important?
Field data is used for aggregations and sorting on certain field types.
Q129: What is a runtime field?
A runtime field is computed at query time rather than stored in the index.
Q130: What is a scripted field?
A scripted field is calculated using scripts during a query.
Q131: Why are scripts often expensive?
Because they run at query time and can be computationally expensive.
Q132: What is a `search` API?
It queries documents from Elasticsearch.
Q133: What is a `count` API?
It counts matching documents.
Q134: What is a `bulk` API?
It allows high-throughput ingestion of many documents in one request.
Q135: What is a `cat` API?
It provides compact human-readable cluster, node, and index information.
Q136: Why do operators use `cat` often?
Because it is convenient for quick cluster diagnostics.
Q137: What is Elasticsearch monitoring?
Monitoring tracks cluster health, shard utilization, memory, and indexing throughput.
Q138: What is a cluster state?
A cluster state contains topology and metadata about nodes, indices, and shards.
Q139: What is a node role?
Node role may be master, data, ingest, coordinating-only, or combinations.
Q140: Why separate node roles?
To improve scalability and isolate workloads.
Q141: What is ingest pipeline?
An ingest pipeline transforms documents before they reach an index.
Q142: What is Logstash filter chain?
A Logstash pipeline can have multiple filters such as grok, mutate, geoip, and date.
Q143: Why add filters before Elasticsearch?
Because it normalizes, enriches, and structures the data before indexing.
Q144: What is a pipeline bug?
It can misparse or corrupt field values, leading to poor indexing and search results.
Q145: What is a no-index pattern?
It means some data is intentionally not stored in Elasticsearch because it is irrelevant or too noisy.
Q146: Why use drop filters?
To discard irrelevant or duplicate events before they reach Elasticsearch.
Q147: What is a log transform?
A transformation reorganizes or enriches raw logs into structured fields.
Q148: What is telemetry?
Telemetry is the collection of operational data like logs, events, metrics, and traces.
Q149: What is log correlation?
It links related log events across services or components for troubleshooting.
Q150: What is a trace ID in logs?
A trace ID allows associating log events with a specific request or transaction.
Q151: Why are trace IDs important?
They aid debugging distributed systems and correlate logs across services.
Q152: What is a multi-line log?
A multi-line log is a single event spread across several lines.
Q153: Why parse multi-line logs carefully?
Because they may otherwise be indexed as multiple unrelated events.
Q154: What is a JSON log?
A JSON log is structured and easy to parse and query.
Q155: Why prefer JSON logs?
They are easier to index and query than plain text logs.
Q156: What is a logstash `json` codec?
It parses JSON payloads directly into structured key-value fields.
Q157: What is a `csv` input?
It reads comma-separated values, often used for structured exports.
Q158: What is a `jdbc` input?
A `jdbc` input reads database rows into Logstash and Elasticsearch.
Q159: What is a `beats` input?
A Beats input collects events from Beats agents such as Filebeat, Metricbeat, and Heartbeat.
Q160: What is a pipeline error?
It may be due to bad field names, parse failures, or syntax issues in Logstash config.
Q161: What is Logstash output to Elasticsearch cluster?
It sends processed documents to one or more Elasticsearch endpoints.
Q162: Why consider data shaping before indexing?
Because fields and mappings are crucial for search quality and aggregation accuracy.
Q163: What is a `dropif` filter?
It drops data matching a condition, often used for noisy or sensitive logs.
Q164: What is a `mutate` filter in Logstash?
It changes fields such as adding, renaming, or converting values.
Q165: What is an event timestamp?
The time associated with the event, often derived from the log or the ingestion time.
Q166: Why is timestamp parsing essential?
Because time-based queries and dashboards depend on a valid event time.
Q167: What is a date filter in Logstash?
It parses dates into Elasticsearch-compatible timestamp values.
Q168: What is a field extraction pipeline?
It turns raw text into structured fields such as service, host, severity, and message.
Q169: What is a data sink?
It is the destination of processed data, often Elasticsearch.
Q170: What is a filter chain bug?
It happens when one filter transforms fields in a way that breaks later processing.
Q171: What is event enrichment?
It adds metadata such as environment, hostname, or region to logs.
Q172: Why enrich logs?
Because it makes them more context-rich and easier to operationalize.
Q173: What is a log signature?
A log signature is a pattern or fingerprint that identifies a recurring event pattern.
Q174: What is log correlation across hosts?
It connects events from multiple systems related to the same transaction or incident.
Q175: What is an ELK dashboard alert?
A dashboard alert identifies anomalies or threshold breaches in cluster data.
Q176: What is a threshold rule?
A threshold rule triggers an alert when a metric or log count exceeds a value.
Q177: What is a log TTL?
TTL (time-to-live) defines how long a document should stay before expiration.
Q178: Why use TTL or retention?
To manage storage cost and lifecycle of the logs.
Q179: What is data archival?
Archival stores old logs in cheaper storage for later or compliance-driven access.
Q180: Why do large enterprises implement hot/warm/cold cycles?
Because search speed, cost, and retention needs differ across data age and access frequency.
Advanced / Expert
Q181: What is Elasticsearch cluster sizing?
Sizing is the process of choosing hardware and shard counts based on expected data volume and query load.
Q182: What is a shard count trade-off?
More shards improve parallelism but increase overhead and management cost.
Q183: Why is shard balancing important for performance?
Because uneven distribution can create hot spots and slow queries.
Q184: What is a heap size tuning issue?
Elasticsearch heap size must be tuned carefully; too high or too low can hurt performance.
Q185: Why is JVM tuning important for Elasticsearch?
Elasticsearch runs on the JVM and is sensitive to heap, GC, and memory pressure.
Q186: What is the index refresh interval trade-off?
A short refresh interval makes search more real-time but increases indexing overhead.
Q187: What is ingestion backpressure?
Backpressure occurs when the pipeline cannot keep up with event rates.
Q188: How do operators handle backpressure?
By buffering, scaling out, rate limiting, or rebalancing data pipelines.
Q189: What is data durability in Elasticsearch?
Durability means data is stored safely and can survive node or shard failures.
Q190: Why are replicas important to durability?
They protect against node failure and can help cluster recovery.
Q191: What is the role of Elasticsearch rebalancing?
It moves shards across nodes to keep the cluster balanced and healthy.
Q192: What is query performance tuning?
It involves good mappings, efficient filters, reduced field data, and proper index design.
Q193: What is search latency?
It is the time taken to complete a search query.
Q194: Why do aggregations matter for operational platforms?
They help answer business and operational questions like “which service failed the most?” or “what is the p99 latency?”
Q195: What is `p95`, `p99`, and percentiles in Elasticsearch?
They describe the latency thresholds for the slowest 5% or 1% of requests.
Q196: Why are percentiles useful?
Because average latency can hide spikes and bad tail behavior.
Q197: What is a hot shard?
A hot shard is a shard receiving disproportionate traffic or load.
Q198: What is a cluster hotspot?
A hotspot is a node or shard receiving much more traffic than its peers.
Q199: What are common operational concerns in ELK?
- cluster health
- shard allocation
- index lifecycle
- disk pressure
- mapping drift
- ingest backlog
- query performance
Q200: What is the main lesson of ELK?
ELK is not just log storage; it is an operational analytics platform that turns raw event streams into searchable, filterable, and visualizable evidence of system health, behavior, and errors.