Spring Kafka/Messaging
Spring Kafka/Messaging
Beginner
Q1: What is messaging in distributed systems?
Asynchronous communication where producers send messages and consumers process them.
Q2: Why use messaging?
Decouples services, improves scalability, and enables asynchronous workflows.
Q3: What is Apache Kafka?
A distributed event streaming platform for high-throughput, durable message processing.
Q4: What is Spring for Apache Kafka?
Spring project providing convenient Kafka producer/consumer integration.
Q5: What is a topic?
Named stream/category where Kafka records are published.
Q6: What is a partition?
Ordered sub-log of a topic enabling parallelism and scalability.
Q7: What is an offset?
Position of a record within a partition log.
Q8: What is a producer?
Client that publishes records to Kafka topics.
Q9: What is a consumer?
Client that reads records from Kafka topics.
Q10: What is a consumer group?
Set of consumers sharing work of topic partitions.
Q11: How are partitions assigned in a group?
Each partition is consumed by at most one consumer in a group at a time.
Q12: What is a broker?
Kafka server node storing and serving topic partitions.
Q13: What is replication factor?
Number of copies of each partition across brokers.
Q14: Why replication matters?
Improves fault tolerance and availability.
Q15: What is leader partition replica?
Replica handling reads/writes for a partition.
Q16: What are follower replicas?
Replicas syncing from leader for redundancy.
Q17: What is ISR (in-sync replicas)?
Replicas sufficiently caught up with leader.
Q18: What is Kafka record structure?
Key, value, headers, timestamp, topic, partition, offset metadata.
Q19: Why use message keys?
Control partitioning and preserve per-key ordering.
Q20: What is ordering guarantee in Kafka?
Order is guaranteed within a partition, not across partitions.
Q21: What is Spring KafkaTemplate?
Main helper for sending messages in Spring applications.
Q22: How do you consume messages in Spring Kafka?
Typically with @KafkaListener.
Q23: What is @KafkaListener?
Annotation marking method to receive Kafka messages from topics.
Q24: What is container factory in Spring Kafka?
Config object controlling listener container behavior.
Q25: What is deserializer?
Converts byte data from Kafka into Java objects.
Q26: What is serializer?
Converts Java objects to bytes for Kafka transmission.
Q27: Common payload formats?
JSON, Avro, Protobuf, plain strings, binary custom formats.
Q28: What is bootstrap servers config?
Initial broker addresses for Kafka client connection discovery.
Q29: What is ack in producer send?
Broker acknowledgment level for write durability behavior.
Q30: Common producer acks values?
0, 1, all (or -1 equivalent).
Q31: What does acks=all imply?
Leader waits for all in-sync replicas before acknowledging.
Q32: What is at-most-once delivery?
Messages may be lost, but never redelivered.
Q33: What is at-least-once delivery?
Messages are not lost (under expected conditions) but may be duplicated.
Q34: What is exactly-once semantics (EOS) concept?
Processing/output appears once logically, requiring transactional/idempotent patterns.
Q35: Is “exactly once” simple in distributed systems?
No, it needs careful end-to-end design.
Q36: What is auto-commit in consumers?
Periodic automatic offset commits by client.
Q37: Why can auto-commit be risky?
Offsets may commit before processing truly succeeds.
Q38: What is manual acknowledgment?
Application explicitly commits offsets after successful processing.
Q39: Why use manual ack?
Better control over processing and failure semantics.
Q40: What is rebalance?
Redistribution of partitions among consumers in a group.
Q41: What can trigger rebalance?
Consumer join/leave, crashes, subscription changes, metadata changes.
Q42: Why are frequent rebalances bad?
Pause processing and increase latency/duplicate risk.
Q43: What is dead-letter topic (DLT)?
Topic where failed/unprocessable messages are routed.
Q44: Why use dead-letter topics?
Preserve problematic events for analysis/replay without blocking pipeline.
Q45: What is retry topic pattern?
Failed messages retried via delayed/retry topics before DLT.
Q46: What is idempotent consumer?
Consumer logic safe against duplicate message processing.
Q47: Why idempotency is critical?
At-least-once delivery means duplicates can happen.
Q48: What is poison pill message?
Record consistently failing processing due to bad data/schema/business rule.
Q49: How to handle poison pills?
Route to DLT with diagnostics instead of infinite retries.
Q50: What is message header in Kafka?
Metadata key-value pairs attached to records.
Q51: Header use cases?
Trace IDs, schema version, tenant, source system identifiers.
Q52: What is lag?
Difference between produced and consumed offsets.
Q53: Why monitor consumer lag?
Indicates backlog growth and processing health.
Q54: What is retention in Kafka?
How long records are kept (time/size policy).
Q55: What is compaction?
Retention mode keeping latest record per key.
Q56: When use compacted topics?
State/event-upsert streams where latest key state matters.
Q57: Beginner messaging anti-pattern?
Assuming no duplicates and writing non-idempotent consumers.
Q58: Beginner observability baseline?
Track throughput, lag, error rates, retry/DLT counts.
Q59: Beginner security baseline for Kafka?
TLS, authentication, ACL authorization, secret management.
Q60: Beginner best practice?
Design for retries, duplicates, and failure from day one.
Intermediate
Q61: What is producer idempotence in Kafka?
Feature ensuring producer retries don’t create duplicates on broker log per session.
Q62: How enable producer idempotence conceptually?
Set producer idempotence configuration and compatible reliability settings.
Q63: What is transactional producer?
Producer that groups writes (and optionally offsets) into atomic transactions.
Q64: What is readcommitted isolation?
Consumers read only committed transactional records.
Q65: What is readuncommitted?
Consumers may read uncommitted/aborted transactional records.
Q66: What is transactional.id?
Producer identity enabling transactional guarantees across restarts/sessions.
Q67: What is outbox pattern with Spring Kafka?
Write domain change + outbox row in DB transaction, publish asynchronously to Kafka.
Q68: Why outbox pattern?
Avoid dual-write inconsistency between DB and Kafka send.
Q69: What is dual-write problem?
DB commit succeeds but publish fails, or publish succeeds but DB rollback.
Q70: CDC vs polling outbox?
CDC streams DB changes; polling periodically reads outbox table.
Q71: What is @Transactional with Kafka integration caveat?
DB and Kafka transactions are distinct unless explicitly coordinated pattern used.
Q72: What is Chained transaction manager concept (historical/complex)?
Attempts multi-resource coordination; often replaced by outbox for simplicity/resilience.
Q73: What is concurrency setting in @KafkaListener container?
Number of concurrent consumer threads/containers.
Q74: How does concurrency relate to partitions?
Effective parallelism limited by partition count per group.
Q75: What if concurrency > partitions?
Extra threads idle for that topic/group combination.
Q76: What is max.poll.records?
Max records returned in one poll call.
Q77: Why tune max.poll.records?
Balance throughput, processing time, and rebalance timeout risk.
Q78: What is max.poll.interval.ms?
Max allowed delay between polls before consumer considered failed.
Q79: Long processing risk with max.poll.interval?
Consumer removed from group causing rebalance/duplicates.
Q80: What is session.timeout.ms?
Failure detection timeout for consumer heartbeats.
Q81: What is heartbeat interval?
Frequency of consumer liveness heartbeats to group coordinator.
Q82: What is pause/resume consumption?
Temporarily stop/resume partition fetching for backpressure/control.
Q83: When pause partitions?
Downstream outage, overload, or controlled throttling scenarios.
Q84: What is CommonErrorHandler (Spring Kafka)?
Centralized listener error handling strategy in modern Spring Kafka.
Q85: What is DefaultErrorHandler?
Configurable error handler supporting retries, backoff, and recovery.
Q86: What is BackOff strategy?
Delay policy between retry attempts (fixed/exponential).
Q87: Why exponential backoff?
Reduces pressure on failing dependencies over time.
Q88: What is DeadLetterPublishingRecoverer?
Publishes failed records to DLT after retries exhausted.
Q89: What context should be added to DLT messages?
Original topic/partition/offset, exception type, stack summary, timestamp, trace ID.
Q90: What is seek in consumer error handling?
Move offset position to reprocess or skip records intentionally.
Q91: What is ack mode in listener containers?
Defines when offsets are committed (record, batch, manual, etc.).
Q92: Record ack mode vs batch ack mode?
Commit per record vs per batch of polled records.
Q93: What is batch listener?
Listener receives list of records for batch processing.
Q94: Batch listener tradeoff?
Higher throughput but more complex error granularity.
Q95: What is record filtering strategy?
Drop duplicates/unwanted records before listener logic.
Q96: What is RecordFilterStrategy?
Spring hook for filtering records pre-processing.
Q97: Why include schema version in messages?
Supports backward/forward-compatible deserialization and migration.
Q98: What is schema registry concept?
Central service storing and validating message schemas.
Q99: Why use Avro/Protobuf with schema registry?
Strong contracts and evolution controls for event payloads.
Q100: What is backward compatibility in schema evolution?
New schema can read old data.
Q101: What is forward compatibility?
Old consumers can read new producer data under defined rules.
Q102: What is full compatibility?
Both backward and forward compatibility constraints satisfied.
Q103: What is tombstone record?
Compacted-topic record with key and null value indicating delete semantics.
Q104: Why careful with null payload handling?
Deserializer/business logic must distinguish tombstones from malformed data.
Q105: What is partition key design principle?
Stable business key balancing ordering needs and partition distribution.
Q106: What is hot partition problem?
Uneven key distribution overloads specific partition/consumer.
Q107: How mitigate hot partitions?
Better key strategy, more partitions, key salting (carefully), workload redesign.
Q108: What is message ordering vs parallelism tradeoff?
More partitions improve parallelism but ordering is per key/partition only.
Q109: What is request-reply messaging over Kafka?
Pattern using reply topics/correlation IDs for pseudo-synchronous workflows.
Q110: When avoid request-reply over Kafka?
Low-latency synchronous requirements often better served by HTTP/gRPC.
Q111: What is exactly-once in Kafka Streams vs consumer apps?
Kafka Streams has built-in EOS features; plain consumers need explicit patterns.
Q112: What is retry topic orchestration pattern?
Route failures through staged retry topics with increasing delays.
Q113: What is non-blocking retry concept?
Main consumer continues processing newer messages while failed ones retry asynchronously.
Q114: What is blocking retry concept?
Consumer thread retries immediately, potentially reducing throughput.
Q115: What is intermediate anti-pattern in messaging?
Infinite retries with no DLT and no observability.
Q116: Better failure strategy?
Bounded retries + backoff + DLT + replay tooling.
Q117: What is replay strategy?
Reprocess historical or DLT events safely with idempotent logic.
Q118: Why replay tooling is essential?
Incidents/data fixes often require controlled reprocessing.
Q119: What is secure deserialization concern?
Avoid unsafe polymorphic deserialization that enables gadget attacks.
Q120: How harden deserialization?
Trusted packages, explicit type mapping, schema validation, minimal polymorphism.
Q121: What is ACL in Kafka?
Authorization rules controlling topic/group operations.
Q122: Why principle of least privilege for Kafka ACLs?
Limit blast radius of compromised credentials/services.
Q123: Intermediate maturity signal?
Team can explain delivery semantics and offset strategy per consumer.
Q124: Intermediate testing strategy?
Embedded/integration tests with real broker containers + failure scenarios.
Q125: Why test rebalance scenarios?
They frequently reveal duplication/order and timeout bugs.
Q126: What is contract testing for events?
Validate producer/consumer schema and semantic expectations continuously.
Q127: Why monitor DLT rate trends?
Spikes indicate upstream regressions or schema/data drift.
Q128: What is lag-based autoscaling concept?
Scale consumers based on lag/throughput metrics.
Q129: Why lag alone can mislead scaling?
Need partition limits, processing cost, and downstream capacity context.
Q130: Intermediate best practice?
Optimize for correctness first, then throughput with measured tuning.
Advanced
Q131: What is end-to-end exactly-once challenge?
Requires coordinated idempotency/transactions across producers, brokers, consumers, and sinks.
Q132: Why Kafka EOS alone may be insufficient?
External DB/API side effects still need idempotent handling.
Q133: What is idempotency key store pattern?
Persist processed event IDs/keys to prevent duplicate side effects.
Q134: Tradeoff of idempotency store?
Extra storage/lookups vs correctness under retries/replays.
Q135: What is transactional outbox relay scaling concern?
Relay throughput, lock contention, and ordering guarantees need careful design.
Q136: Polling outbox vs CDC advanced tradeoff?
Polling simpler but adds query load/latency; CDC near-real-time but operationally heavier.
Q137: What is event versioning strategy?
Use explicit version fields and compatibility rules per event type.
Q138: What is schema evolution anti-pattern?
Breaking field changes without staged consumer migration.
Q139: What is consumer-driven event evolution?
Producers coordinate changes based on downstream consumer compatibility needs.
Q140: What is event envelope pattern?
Standard metadata wrapper (id, type, version, timestamp, trace) around payload.
Q141: Why event envelope helps?
Consistent observability, routing, governance across topics.
Q142: What is data lineage in event systems?
Tracking origin and transformations of data across services/topics.
Q143: What is poison message quarantine workflow?
DLT + triage + fix + controlled replay + postmortem.
Q144: What is deterministic replay requirement?
Reprocessing should produce predictable results with versioned logic.
Q145: What is ordering recovery challenge after retries?
Retry paths can alter apparent processing order across partitions/topics.
Q146: How preserve per-key ordering under retries?
Keyed partitions + careful retry routing + ordered consumer execution policies.
Q147: What is throughput collapse during incidents?
Retries and downstream slowness consume consumer capacity.
Q148: How prevent retry storms in Kafka consumers?
Bound retries, circuit breakers, adaptive pause, retry budgets.
Q149: What is backpressure strategy in message consumers?
Control poll/processing rate and downstream concurrency limits.
Q150: What is cooperative rebalancing concept?
Rebalance strategy reducing stop-the-world partition movement impact.
Q151: Why static membership can help?
Reduces disruptive rebalances during brief restarts.
Q152: What is rack awareness in Kafka?
Replica placement across failure domains for resilience.
Q153: What is unclean leader election risk?
Can improve availability but risks data loss.
Q154: What is min.insync.replicas importance?
Defines durability threshold for successful writes with stronger ack settings.
Q155: acks=all without proper min.insync.replicas risk?
False sense of durability if ISR requirement too low.
Q156: What is compression in Kafka producers?
Compress batches (snappy/lz4/zstd/gzip) to reduce bandwidth/storage.
Q157: Compression tradeoff?
CPU overhead vs network/storage efficiency.
Q158: What is linger.ms effect?
Wait briefly to batch more records, improving throughput with slight latency cost.
Q159: What is batch.size effect?
Producer batch memory target influencing throughput/latency profile.
Q160: What is max.in.flight.requests.per.connection concern?
Affects ordering guarantees during retries for some configurations.
Q161: What is advanced security baseline for Kafka platforms?
TLS+mTLS, SCRAM/OAuth auth, ACL governance, secret rotation, audit logs.
Q162: What is message-level encryption need?
Protect sensitive payload fields end-to-end beyond transport encryption.
Q163: What is PII governance in event streams?
Schema classification, minimization, retention controls, deletion workflows.
Q164: What is right-to-erasure challenge with Kafka?
Immutable logs complicate deletion; use compaction/tombstones and data minimization strategies.
Q165: What is multi-cluster Kafka strategy?
Replication/failover across clusters for DR and geo requirements.
Q166: What is active-active messaging complexity?
Conflict resolution, ordering, duplication, and failover coordination challenges.
Q167: What is MirrorMaker/replication role conceptually?
Replicate topics across clusters for disaster recovery/regional distribution.
Q168: What is SLA/SLO for messaging systems?
Targets for publish latency, consumer lag, durability, and availability.
Q169: What metrics are critical for Spring Kafka consumers?
Lag, poll duration, processing latency, retry counts, DLT volume, rebalance frequency.
Q170: What metrics are critical for producers?
Send latency, error rate, record retry rate, batch utilization, buffer exhaustion.
Q171: What is tracing strategy for async messaging?
Propagate trace context in headers and create spans for produce/consume/process.
Q172: What is incident response playbook for Kafka failures?
Identify scope, isolate failing consumers, control retries, drain/replay safely.
Q173: What is game day for messaging systems?
Planned drills for broker outages, lag spikes, schema breakages, replay operations.
Q174: Biggest advanced Spring Kafka anti-pattern?
Treating messaging as fire-and-forget without contracts, observability, and replay design.
Q175: What is mature event-driven architecture indicator?
Clear ownership of schemas, semantics, retries, and recovery workflows.
Q176: Final reliability principle for messaging?
Assume duplicates, delays, and partial failures are normal conditions.
Q177: Final performance principle?
Tune producers/consumers empirically with production-like workloads.
Q178: Final security principle?
Secure brokers, clients, payloads, and operational access paths end-to-end.
Q179: Final operations principle?
Automate monitoring, DLT handling, replay tooling, and runbooks.
Q180: Final maturity principle?
Spring Kafka excellence means predictable correctness and operability at scale.