Kafka operations require evidence across brokers/controllers, producers, consumers, stream processors, networks, storage, and downstream systems. One dashboard or consumer-lag number cannot establish health.

Define service objectives

Set availability, durability, end-to-end freshness, throughput, p95/p99 latency, recovery time, recovery point, and cost objectives by workload. Record partition ordering, retention, replay, and data-loss tolerances.

Monitor layers

LayerSignals
ClusterController/quorum health, under-replicated/offline partitions, ISR changes, disk, network, request latency
ProducerErrors, retries, record latency, batching, compression, throttling
ConsumerLag with input rate, poll/commit errors, rebalances, processing freshness
StreamsTask state, restoration, dropped/late records, store and processing metrics
OutcomeReconciliation, missing/duplicate events, downstream availability and freshness

Test capacity and failure

Use representative record sizes, keys, partitions, compression, replication, retention, consumer groups, and bursts. Exercise broker/controller loss, disk pressure, network partition, zone loss, rebalance, rolling upgrade, credential rotation, throttling, and restore.

Clarify cloud responsibility

Managed services can operate infrastructure while customers retain responsibility for data classification, identity, network access, topic/ACL design, clients, schemas, retention, monitoring, and application recovery. Validate provider-specific availability, quotas, encryption, logs, regions, backup/export, egress cost, and exit plan.

Connect Streams monitoring, advanced configuration, and security and multi-cluster planning.

Reviewed against current Apache Kafka documentation September 4, 2026. Original publication date preserved.

Completion guide. A healthy broker is not the same thing as a healthy event-driven product. Monitor from the producer request through topic durability and consumer progress to the business outcome the event was supposed to create.

The observability signal path

Kafka signal path
01Producer
02Broker & partitions
03Consumer group
04Business outcome
The moving marker represents an event and its operational evidence crossing each boundary. Motion pauses automatically when reduced motion is preferred.

Instrument every boundary with rate, errors, and duration, then add Kafka-specific saturation and correctness signals. End-to-end freshness—a timestamp or synthetic event observed at the destination—is the strongest check that the whole path is working.

Map symptoms to Kafka signals

User-visible symptomKafka evidenceCorrelate withFirst question
Events arrive lateConsumer lag by partition; fetch and processing ratesIngress rate, handler latency, rebalancesIs lag growing everywhere or on one partition?
Writes are slowProduce request latency, timeouts, retriesISR changes, disk/network saturationDid acknowledgement wait or broker service time grow?
Missing resultsOffset movement and error/DLQ countsSchema failures, transaction aborts, sink errorsWas the record never produced, not consumed, or rejected?
Unstable throughputRebalance frequency, partition skewDeployments, GC, autoscaling, hot keysDoes instability align with membership changes?

A runbook for rising consumer lag

  1. Scope. Compare partitions. One hot partition suggests key skew; broad growth suggests insufficient capacity or downstream slowness.
  2. Rate. Compare arrival rate with consume and processing rates. Lag alone has no time dimension.
  3. Stability. Check group membership and rebalance activity around deploys, autoscaling, or timeouts.
  4. Handler. Inspect external calls, retries, GC pauses, and record-size outliers inside the consumer.
  5. Recoverability. Estimate drain time at safe capacity before adding workers; partitions cap useful parallelism.

Failure drill and release gate

Drill 01

Broker interruption

Confirm leaders recover, producers retry within their delivery contract, ISR returns to normal, and no silent loss is observed.

Drill 02

Poison record

Verify deserialization and business-rule failures follow a bounded retry and quarantine policy with enough context to replay safely.

Drill 03

Slow dependency

Throttle a sink and validate backpressure, lag alerts, resource ceilings, and drain-time estimates.

Alert on customer-impacting conditions and sustained trends, not every transient metric movement. Each alert should link to a runbook, name an owner, and state the safe first action.

Primary references and next reading

Use version-matched documentation for configuration defaults. The links below point to maintained upstream documentation rather than copied defaults.