Elasticsearch & Kibana

Elasticsearch & Kibana


Beginner

Q1: What is Elasticsearch?

Elasticsearch is a distributed search and analytics engine built on Apache Lucene.

Q2: What is Kibana?

Kibana is a visualization and management UI for Elasticsearch data.

Q3: Elasticsearch + Kibana in one line?

Elasticsearch stores/searches data; Kibana explores/visualizes/manages it.

Q4: What is an index in Elasticsearch?

Logical namespace containing related documents.

Q5: What is a document?

JSON object stored and indexed in Elasticsearch.

Q6: What is a field?

Key in a document with a typed value used for search/aggregation.

Q7: What is mapping?

Schema definition describing field types and indexing behavior.

Q8: Why mappings matter?

They control correctness, query behavior, and storage/performance characteristics.

Q9: What is dynamic mapping?

Automatic field type inference when new fields appear.

Q10: Dynamic mapping risk?

Mapping explosion and incorrect inferred types.

Q11: What is a shard?

Partition of an index (Lucene index unit).

Q12: What is primary shard?

Original shard holding authoritative indexed data segment.

Q13: What is replica shard?

Copy of primary shard for high availability and read scaling.

Q14: Why replicas are useful?

Fault tolerance and increased search throughput.

Q15: What is cluster in Elasticsearch?

Group of nodes working together as one logical system.

Q16: What is node?

Single Elasticsearch server instance in a cluster.

Q17: What is cluster health status?

Green/yellow/red indicating shard allocation state.

Q18: Green vs yellow vs red?

Green: all shards allocated; yellow: replicas unassigned; red: primaries unassigned.

Q19: What is inverted index?

Data structure mapping terms to documents for fast full-text search.

Q20: What is analyzer?

Pipeline that transforms text into searchable tokens.

Q21: Analyzer components?

Character filters, tokenizer, token filters.

Q22: What is tokenization?

Splitting text into terms/tokens during analysis.

Q23: What is keyword field type?

Exact value field (not analyzed) for filtering/sorting/aggregations.

Q24: Text vs keyword quick rule?

text for full-text search; keyword for exact matching and aggregations.

Q25: What is relevance score?

Numeric ranking value indicating document match quality.

Q26: What is Query DSL?

JSON-based query language for Elasticsearch searches.

Q27: What is match query?

Full-text query using analyzer semantics.

Q28: What is term query?

Exact term lookup (commonly for keyword/numeric fields).

Q29: What is bool query?

Combines must/should/filter/mustnot clauses.

Q30: Filter vs must difference?

Filter is non-scoring and cache-friendly; must affects scoring.

Q31: What is aggregation?

Analytics computation over documents (counts, stats, buckets).

Q32: Common aggregation example?

Terms aggregation for top values by frequency.

Q33: What is range query?

Find docs where field values fall within bounds.

Q34: What is wildcard query?

Pattern-based term matching (can be expensive).

Q35: What is pagination in Elasticsearch?

Returning result subsets via from/size or searchafter.

Q36: What is from/size pitfall?

Deep pagination cost grows significantly.

Q37: What is Kibana Discover?

Interface for ad-hoc document exploration and filtering.

Q38: What is Kibana dashboard?

Collection of visualizations for monitoring/analysis.

Q39: What is Kibana visualization?

Chart/table/metric built from Elasticsearch queries/aggregations.

Q40: What is data view (index pattern concept)?

Kibana definition selecting indices/fields for analysis.

Q41: What is time filter in Kibana?

Global time range constraint for time-based data exploration.

Q42: Why time filtering is important?

Narrows search scope for speed and relevance.

Q43: What is ingestion pipeline concept?

Preprocessing documents before indexing (enrich/transform/parse).

Q44: What is Logstash relation?

Pipeline tool often used to ingest/transform data into Elasticsearch.

Q45: Beats relation to Elasticsearch?

Lightweight shippers (Filebeat/Metricbeat/etc.) for data collection.

Q46: What is beginner anti-pattern in Elasticsearch?

Using default mappings for all fields without planning.

Q47: Another beginner anti-pattern?

Storing extremely high-cardinality fields without purpose.

Q48: Beginner reliability baseline?

Replicas enabled and basic snapshot strategy.

Q49: Beginner performance baseline?

Right field types and controlled shard count.

Q50: Beginner security baseline?

Authentication, TLS, and role-based access enabled.

Q51: What is snapshot in Elasticsearch?

Backup mechanism storing index data/metadata in repository.

Q52: Why snapshots matter?

Primary disaster recovery method for clusters.

Q53: What is reindex API?

Copy/transform documents from one index to another.

Q54: Why reindex is common?

Schema evolution and index migration tasks.

Q55: What is alias in Elasticsearch?

Logical name pointing to one or more indices.

Q56: Alias benefit?

Decouple application index name from physical index versions.

Q57: What is rollover concept?

Create new write index when size/age/doc thresholds reached.

Q58: Why rollover helps?

Manage index growth and retention efficiently.

Q59: Beginner workflow principle?

Design mappings and lifecycle before scaling ingestion.

Q60: Beginner best practice?

Treat search schema as a product contract.

Intermediate

Q61: What is multi-field mapping?

Index same field in multiple ways (e.g., text + keyword).

Q62: Why multi-fields are useful?

Support both full-text and exact-match aggregations on same data.

Q63: What is custom analyzer use case?

Language-specific stemming, synonyms, edge n-grams, normalization.

Q64: What is normalizer?

Analyzer-like processing for keyword fields (no tokenization).

Q65: What is synonym filter?

Expands equivalent terms during analysis or query time.

Q66: Synonym management pitfall?

Incorrect expansion can reduce precision and relevance quality.

Q67: What is stemming?

Reducing words to root forms for broader matching.

Q68: What is fuzziness in match queries?

Tolerance for edit distance typos.

Q69: Fuzzy query tradeoff?

Better recall vs potentially higher query cost/noise.

Q70: What is minimumshouldmatch?

Controls required should-clause match proportion/count.

Q71: What is boosting?

Increasing influence of fields/clauses on relevance ranking.

Q72: What is functionscore query?

Custom scoring based on numeric signals/functions.

Q73: What is nested field type?

Model arrays of objects preserving per-object field relationships.

Q74: Why nested queries needed?

Avoid cross-object matching errors in arrays of objects.

Q75: What is parent-child join?

Relationship model across documents in same index (specialized use).

Q76: Parent-child tradeoff?

Flexibility vs query/indexing complexity and overhead.

Q77: What is runtime field?

Field computed at query time without full reindex.

Q78: Runtime field tradeoff?

Flexibility vs higher query-time cost.

Q79: What is docvalues?

Columnar storage for sorting/aggregations/script access.

Q80: Why docvalues matter?

Efficient aggregations and sorting on disk-backed structures.

Q81: What is fielddata?

Heap-based data structure for text aggregations/sorts (costly).

Q82: Why avoid fielddata on text generally?

High memory usage; prefer keyword multi-fields.

Q83: What is refresh interval?

Frequency making indexed docs searchable (near real-time behavior).

Q84: Lower refresh interval tradeoff?

Faster visibility vs higher indexing overhead.

Q85: What is translog?

Write-ahead transaction log for durability/recovery.

Q86: What is segment in Lucene context?

Immutable index file unit created during indexing.

Q87: What is merge process?

Background compaction of segments for efficiency.

Q88: Merge impact concern?

I/O and CPU spikes affecting query/index performance.

Q89: What is index template?

Reusable settings/mappings/aliases for new indices or data streams.

Q90: Component template?

Composable reusable part of index template.

Q91: What is ILM?

Index Lifecycle Management automating rollover, warm/cold/delete phases.

Q92: Why ILM is important?

Cost control and retention governance at scale.

Q93: What is data stream?

Time-series abstraction managing backing indices automatically.

Q94: Data stream vs classic index?

Optimized write-once append-heavy time-series lifecycle workflows.

Q95: What is ingest pipeline processor?

Built-in transform step (grok, date, geoip, set, remove, etc.).

Q96: What is grok processor?

Pattern-based text parsing into structured fields.

Q97: What is painless scripting?

Elasticsearch scripting language for updates, scoring, transforms.

Q98: Scripting risk?

Performance and security implications if overused/unbounded.

Q99: What is circuit breaker in Elasticsearch?

Protective limit preventing out-of-memory from expensive operations.

Q100: What is query cache?

Cache for filter context results to accelerate repeated queries.

Q101: What is request cache?

Cache for aggregation-heavy identical requests (conditions apply).

Q102: What is searchafter?

Efficient deep pagination using sort values from previous page.

Q103: searchafter vs scroll?

searchafter for user pagination; scroll for large batch extraction/reprocessing.

Q104: What is point-in-time (PIT)?

Consistent search snapshot context for paginated queries.

Q105: What is slowlog?

Logs slow indexing/search operations for tuning diagnostics.

Q106: What is hot-warm-cold architecture?

Tiered node roles/storage classes by data age/access frequency.

Q107: Why data tiers help?

Optimize cost-performance across retention lifecycle.

Q108: What is intermediate anti-pattern?

Too many tiny shards causing overhead and poor performance.

Q109: Better shard strategy?

Right-size shard count based on data volume and query patterns.

Q110: What is shard rebalancing?

Redistributing shards across nodes for balance/resilience.

Q111: What is allocation awareness?

Shard placement respecting zones/racks for fault tolerance.

Q112: What is Kibana Lens?

User-friendly drag-and-drop visualization builder.

Q113: What is Kibana alerting concept?

Rule engine for threshold/query/anomaly notifications.

Q114: What is watcher/alert action use?

Send notifications/webhooks/tickets on conditions.

Q115: What is RBAC in Elastic Stack?

Role-based permissions for indices, features, spaces.

Q116: What are Kibana spaces?

Logical UI separation for dashboards/saved objects/access control.

Q117: Intermediate observability baseline?

Track query latency, indexing throughput, JVM heap, shard states.

Q118: Intermediate security baseline?

TLS everywhere, least-privilege roles, audit logging.

Q119: Intermediate reliability baseline?

Snapshot automation and restore drills.

Q120: Intermediate governance baseline?

Template/mapping standards and lifecycle policies.

Q121: Intermediate maturity signal?

Team can evolve mappings with minimal downtime.

Q122: Intermediate ops principle?

Benchmark before and after major schema/query changes.

Q123: Intermediate cost principle?

Use ILM tiers and prune unused fields aggressively.

Q124: Intermediate architecture principle?

Separate ingest-heavy and query-heavy workloads when needed.

Q125: Intermediate best practice?

Design for predictable performance under realistic load.

Advanced

Q126: What is cluster coordination role (master-eligible nodes)?

Manage cluster state updates and shard allocation decisions.

Q127: Why dedicated master nodes?

Improve cluster stability by isolating coordination workload.

Q128: What is split-brain risk (historical concept)?

Cluster partition leading to conflicting masters/state (mitigated by modern coordination).

Q129: What is quorum relevance in cluster state?

Majority decisions protect consistency during failures.

Q130: What is cluster state size challenge?

Large mappings/indices can slow state publication and operations.

Q131: Cluster state optimization tactics?

Template discipline, field limits, index count control.

Q132: What is mapping explosion?

Excessive unique fields causing memory/state/performance problems.

Q133: Mapping explosion mitigation?

dynamic: strict/false, field caps, ingestion normalization.

Q134: What is index sorting?

Pre-sorting segments by fields to accelerate certain queries.

Q135: Index sorting tradeoff?

Faster query patterns vs indexing overhead and rigidity.

Q136: What is kNN/vector search in Elasticsearch?

Approximate nearest-neighbor search over dense vector embeddings.

Q137: Hybrid search concept?

Combine lexical BM25 and vector semantic retrieval.

Q138: Why hybrid search is useful?

Balances precision of keyword with semantic recall.

Q139: What is relevance tuning lifecycle?

Offline judgments + online metrics + iterative analyzer/query adjustments.

Q140: What is search quality metric example?

NDCG, precision@k, recall@k (evaluation-dependent).

Q141: What is cross-cluster search (CCS)?

Query multiple clusters from one coordinating endpoint.

Q142: What is cross-cluster replication (CCR)?

Replicate indices between clusters for DR/locality.

Q143: CCR use cases?

Disaster recovery, geo-distributed read locality.

Q144: What is snapshot lifecycle management (SLM)?

Automates snapshot scheduling and retention.

Q145: Why SLM + ILM together?

Coordinate backup and retention across data lifecycle.

Q146: What is zero-downtime reindex pattern?

Create new index version, reindex, swap alias atomically.

Q147: Why alias swap is powerful?

Instant cutover with rollback option.

Q148: What is write alias pattern?

Stable logical write target pointing to current active index.

Q149: What is ingestion backpressure concern?

Input rate exceeds indexing capacity causing lag/failures.

Q150: Backpressure mitigation?

Queue buffering, bulk tuning, scaling ingest nodes, pipeline optimization.

Q151: What is bulk API?

Efficient batched indexing/update/delete request format.

Q152: Bulk sizing tradeoff?

Too small wastes overhead; too large increases memory/latency risk.

Q153: What is refresh/replica tuning during bulk loads?

Temporarily reduce refresh frequency/replicas for faster ingestion (with risk controls).

Q154: What is advanced cache pitfall?

Low-selectivity/high-churn queries reduce cache effectiveness.

Q155: What is query profiling?

Detailed breakdown of query execution phases for optimization.

Q156: What is security hardening baseline at scale?

mTLS, fine-grained roles, encrypted snapshots, network segmentation.

Q157: What is field-level/document-level security?

Restrict access to specific fields/docs by roles/policies.

Q158: What is compliance concern for logs/search data?

PII retention, right-to-delete, auditability, regional regulations.

Q159: Data minimization strategy?

Index only required fields and redact sensitive content at ingest.

Q160: What is noisy neighbor issue in shared clusters?

One tenant’s heavy indexing/query load degrades others.

Q161: Mitigation for noisy neighbors?

Separate clusters, quotas, index-level controls, workload isolation.

Q162: What is cost governance for Elastic?

Lifecycle tiers, frozen/cold storage, shard optimization, query budgets.

Q163: What is Kibana scale challenge?

Dashboard/query sprawl causing backend load spikes.

Q164: Kibana governance tactics?

Review dashboards, enforce ownership, archive unused assets.

Q165: What is incident response with Elastic stack?

Use logs/metrics/traces correlation and saved investigations/runbooks.

Q166: What is observability convergence pattern?

Combine Elasticsearch logs with metrics/traces for triage context.

Q167: What is advanced anti-pattern?

Treating Elasticsearch as primary OLTP database for heavy transactions.

Q168: Better architecture principle?

Use Elasticsearch for search/analytics; keep source of truth in transactional store.

Q169: What is DR strategy for Elastic clusters?

Snapshots, tested restore, CCR where needed, infra-as-code rebuild paths.

Q170: Why test restore regularly?

Backup confidence requires proven recovery execution.

Q171: What is final reliability principle?

Plan shard/layout/lifecycle for failure scenarios, not ideal conditions.

Q172: What is final performance principle?

Tune mappings, queries, and shard strategy with production-like benchmarks.

Q173: What is final security principle?

Enforce least privilege and encryption across ingest, storage, and access.

Q174: What is final governance principle?

Schema and dashboard changes must be reviewed/versioned like code.

Q175: What is final cost principle?

Retention and tiering should reflect actual query value.

Q176: What is final scaling principle?

Scale horizontally with clear node roles and workload separation.

Q177: What is final search-quality principle?

Continuously evaluate relevance with real user behavior and test sets.

Q178: What is final operations principle?

Instrument cluster health and automate runbooks for common failures.

Q179: What is final collaboration principle?

Platform, data, and app teams co-own schema and search outcomes.

Q180: Final maturity principle?

Elasticsearch/Kibana excellence is secure, relevant, and operationally disciplined search at scale.

Bonus: Minimal Index Template Example (Conceptual)

{
  "index_patterns": ["logs-app-*"],
  "template": {
    "settings": {
      "number_of_shards": 3,
      "number_of_replicas": 1
    },
    "mappings": {
      "dynamic": "strict",
      "properties": {
        "@timestamp": { "type": "date" },
        "level": { "type": "keyword" },
        "message": { "type": "text" },
        "service": { "type": "keyword" },
        "duration_ms": { "type": "long" }
      }
    }
  }
}