FinOps Inform

Platform Owners: Bill First Kafka Cost Optimization With Verified Pilots

Bill first playbook for CTOs and platform owners: measure egress, storage, and connector costs, apply six low risk Kafka changes, then run a verified pilot.

Anonymous Kafka broker storage in data centre

Cut the largest share of your Kafka bill by attacking cross-AZ transfer, retained bytes and compression settings, then right-size brokers before you touch architecture. Start by pulling your transfer, storage and connector worker costs from the last billing cycle: that single report tells you which lever to pull first.


TL;DR:

  • Cross-availability zone data transfer and replicated storage are the main cost drivers, often outweighing broker compute expenses.
  • Implementing producer compression, batching, and rack awareness can reduce costs significantly with minimal operational risk.
  • Tiered and diskless storage models can sharply lower EBS and inter-AZ replication expenses, but may introduce latency and API call costs.
  • Regular cost audits, automation, and ownership attribution are essential for maintaining savings and preventing cost creep over time.
  • Forecasting should focus on data volume, retention decisions, and workload patterns, revisited whenever storage or throughput changes occur.

Where Kafka money actually goes: cost drivers you must measure

Before you change a single config value, you need to know where the money is going. Kafka bills are rarely dominated by compute the way most engineers assume. Cross-AZ data transfer and replicated storage are often the largest line items on a self-managed AWS deployment, with broker compute making up a much smaller share than most teams expect, according to Google Cloud's guidance on managed Kafka costs.

Pull these numbers before you plan anything:

  • Per-AZ egress from your cloud billing console, broken down by source and destination availability zone.
  • GB-month storage costs for each broker's attached volumes, including replica copies.
  • Connector worker hours and their associated compute cost, separated from broker compute.
  • Topic-level retention bytes multiplied by replication factor, which is the real storage footprint you are paying for.
  • Downstream egress to consumers outside the cluster's network, including cross-region reads.

The multiplication that catches teams out is replication factor times retention times consumer fan-out. A topic with a replication factor of three, seven days of retention and five consumer groups reading across availability zones is not one cost centre. They are three physical copies generating cross-AZ traffic on every read, and every extra day of retention multiplies that storage bill by three before you have added a single new consumer.

Configuration levers engineers can change safely (low to medium risk)

Configuration changes carry the lowest operational risk and the fastest payback, which is why they belong first on your list.

  1. Enable producer compression. Test lz4 or zstd against your current setup: zstd typically gives the best compression ratio, but check CPU headroom on producers first, because compression becomes counterproductive once CPU cost exceeds the transfer savings.
  2. Tune batching. Raising linger.ms and batch.size together increases batch efficiency and reduces the number of requests per megabyte sent, which lowers both compute and network cost.
  3. Enable follower fetching and rack awareness. Setting client.rack on consumers and rack IDs on brokers lets consumers read from a same-AZ replica instead of crossing zones, which is one of the more direct fixes for cross-AZ egress, per AWS's guidance on right-sizing Kafka clusters.
  4. Set topic-level retention overrides. Kafka's topic configuration options let you override retention, segment size and min.insync.replicas per topic, so non-critical topics do not need the same retention window as your audit log.
  5. Right-size brokers before you add more. Work backwards from throughput and retention SLOs, load test, and prefer scaling up existing brokers over scaling out when the bottleneck is disk or memory rather than partition count.
  6. Consolidate partitions and review replication factor cautiously, only where durability requirements genuinely allow it.

Pro Tip: Roll out compression and batching changes on one non-critical topic first, measure end-to-end latency for 48 hours, then expand.

Architecture levers: tiered storage and diskless approaches

Bigger gains usually require an architecture shift, and that shift is tiered or diskless storage.

Tiered storage keeps a hot local tier for recent segments and moves cold segments to a remote object store, which reduces broker disk requirements and the cost of rebalancing when brokers fail or scale, according to KIP-405's design description. Diskless, object-backed broker models go further: writes land directly in object storage rather than on attached disks replicated across zones, which can remove cross-AZ replication traffic and much of the EBS bill entirely.

Tiered and diskless storage can sharply cut EBS and inter-AZ replication costs, though the exact saving depends on retention length and read patterns, per KIP-405.

The limitations are real:

  • Remote fetches carry latency that local disk reads do not.
  • Object storage LIST and GET API calls add a cost line that did not exist before.
  • Workloads with heavy historical replay or many small reads against cold segments can erode the savings.

Before piloting, define your rollback criteria, a latency ceiling for remote reads, and the topics best suited to cold storage: typically low-throughput, long-retention topics rather than your hottest streams.

Operational practices and a FinOps playbook for continuous savings

A one-off optimisation pass decays. Costs creep back within months unless you build the audit into your operating rhythm. Stack Overflow's engineering blog frames this correctly: cost-efficient Kafka is an ongoing process, not a project with an end date.

Run these checks on a schedule, not just when a bill spikes:

  • Audit for inactive topics and idle consumer groups quarterly, then delete, archive to object storage, or shorten retention.
  • Automate lifecycle rules for non-production retention and object storage tiers so nobody has to remember to clean up.
  • Instrument per-topic egress and storage cost, and set alerts on anomalies rather than waiting for the monthly invoice.
  • Add cost checks to your PR pipeline for any change that touches retention, replication factor or partition count.

Pro Tip: Assign topic-level cost ownership to the team that produces the data: attribution changes behaviour faster than any dashboard.

Chargebacks and FinOps gating work because they move the cost conversation upstream, into design reviews, instead of downstream into a finance team's spreadsheet nobody in engineering reads.

How Koritsu operationalises Kafka savings

We run a four-step workflow: automated analysis of your bills and Kafka telemetry, root-cause identification tied to specific topics and brokers, hands-on configuration and infrastructure changes, and verification against the actual invoice. An AI agent continuously surfaces the patterns above: idle connectors still running, cross-AZ egress hotspots on specific topics, and retention settings that are candidates for tiering. Specialists then implement the fix rather than leaving it as a recommendation in a report nobody actions. A pilot typically starts with a free assessment that becomes a Savings Opportunity Report, scoped to your actual bill rather than industry averages.

Budget allocation and forecasting for Kafka infrastructure costs

Forecasting Kafka spend badly is easy, because the cost is driven by data volume and retention decisions that shift as products change, not by a fixed headcount or licence fee. Build your forecast from the same components you measured earlier: projected storage growth at current retention settings, expected cross-AZ traffic as new consumers are added, and connector worker hours for planned integrations.

Allocate budget by workload pattern rather than as one lump sum. Steady-state topics with predictable throughput need a stable baseline allocation. Peak workloads, seasonal spikes or event-driven bursts need headroom modelled separately, because provisioning permanent capacity for a rare peak wastes money every day the peak is not happening. Autoscaling connector workers and using burst capacity for object storage tiers handles this better than static overprovisioning.

Revisit the forecast whenever retention policy changes, a new consumer group is onboarded, or a topic's throughput profile shifts meaningfully. A forecast built once at cluster launch and never revisited is the most common reason budgets drift from actuals over a year. Tie the forecast to the same per-topic cost attribution you use for chargebacks, so the team requesting a new topic sees the projected cost before it ships, not after the invoice arrives.

Budget allocation and forecasting for Kafka infrastructure costs โ€” overview diagram

Best practices for cost-effective Kafka security configurations

Security and cost pull in the same direction more often than teams assume, provided you configure deliberately rather than defaulting to the heaviest setting everywhere. TLS encryption on all client and inter-broker traffic adds CPU overhead on brokers, and that overhead scales with throughput, so factor it into your right-sizing calculations rather than discovering it after brokers are underprovisioned.

Authentication and authorisation choices affect operational cost indirectly. SASL/SCRAM is lighter to operate than a full mutual TLS certificate rotation pipeline for smaller teams, but mutual TLS avoids the ongoing cost of credential rotation infrastructure at larger scale. Choose based on your team's existing certificate management maturity rather than defaulting to the most complex option.

Access control lists scoped tightly to topic and consumer group prevent the kind of unauthorised or forgotten access that leads to idle consumers and unused connectors quietly running up connector worker costs. Broad, permissive ACLs make it harder to spot the inactive resources that your audits are trying to catch.

Encryption at rest on tiered or object-backed storage adds a small, predictable cost per gigabyte, which is worth accepting given the storage savings tiering already delivers. The mistake to avoid is applying maximum encryption and audit logging uniformly across every topic regardless of sensitivity, which adds compute and storage cost to data that never needed it.

Best practices for cost-effective Kafka security configurations โ€” overview diagram

Caveats: when optimisation is the wrong call

Do not cut replication, move to single-AZ, or shorten retention where a business SLO or disaster recovery requirement depends on them. Aggressive compression and remote reads add latency: test with canaries and clear rollback triggers before rolling out broadly. Every change in this article should ship with a metric to watch and a plan to reverse it.

Next steps: request a free assessment or pilot with Koritsu

Everything above tells you where to look. Turning that into a verified number on your invoice is the harder part, and it is what we spend our time on. Koritsu starts with a free assessment that becomes a Savings Opportunity Report, scoped to your Kafka clusters and the cost drivers covered here. We only charge a success fee on savings we actually deliver, verified against your bill, and teams that want ongoing coverage can move onto a FinOps as a Service subscription afterwards.

Koritsu AI

If you already know roughly where your Kafka spend is heaviest, check our pricing and success-fee model and get a scoped assessment started this week.

Sources

FAQ

Do Netflix use Kafka?

Netflix is widely known to run Kafka at large scale for event streaming and data pipeline workloads across its infrastructure. Specific architecture details vary by team and change over time, so treat any single description as a snapshot rather than current fact.

Why is LinkedIn replacing Kafka?

LinkedIn originally created Kafka and continues to run it extensively rather than replacing it outright. Reports of specific internal architecture changes at LinkedIn reflect evolving internal tooling choices, not an industry move away from Kafka itself.

Why is ZooKeeper deprecated?

Kafka moved away from ZooKeeper toward KRaft, its own built-in consensus protocol, to remove an external dependency and simplify cluster operations. This shift affects operational complexity more directly than cost, though fewer moving parts generally means fewer components to size and monitor.

How can I make a Kafka consumer faster?

Increase fetch.min.bytes and max.poll.records to pull larger batches per request, and enable follower fetching so consumers read from a same-AZ replica rather than crossing zones. Tuning consumer thread count to match partition count and ensuring deserialization is not a bottleneck also improve throughput without adding brokers.

What typically drives the biggest Kafka cost savings?

Cross-AZ data transfer and replicated storage are usually the largest cost drivers, and configuration changes like compression, batching and follower fetching can typically deliver 20 to 40% in savings. Architecture shifts such as tiered or diskless storage can go further for workloads with long retention and cold-read patterns.