FinOps Inform
FinOps: 90 Days to Size GCP Committed Use Discounts, Kori Report
FinOps teams: size, buy, and monitor GCP committed use discounts using 90 days of billing data. Request a free Kori Savings Opportunity Report to spot...
Committed use discounts let you trade a one or three year usage commitment for deep discounts on eligible Google Cloud spend. Choose resource-based CUDs when your infrastructure footprint is stable and predictable, and spend-based or flexible commitments when your workloads shift shape month to month. The teams that should look at this first are anyone running steady-state Compute Engine, GKE, or Cloud Run workloads for more than a few months at a time.
TL;DR:
- Resource-based commitments offer discounts of up to 55 percent but are limited to specific machine families and region support, with flexibility depending on workload stability.
- Flexible Savings Plans provide monthly entitlement without rollover, making them ideal for unpredictable workloads with seasonal spikes or rapid scaling changes.
- Proper sizing relies on at least ninety days of billing data to identify steady-state usage, with ongoing review essential to avoid stranded costs from workload migration or deprecation.
- Commitments purchased at the billing account level and with clear governance, tagging, and staggered renewals best prevent drifting and wasted discounts.
- Using AI-driven tools to analyze billing data helps teams optimize commitment types and sizes, reducing the risk of overcommitting or underutilizing cloud resources.
What are GCP committed use discounts and which type fits your workload?
Google Cloud offers committed use discounts in exchange for a promise: commit to a minimum level of resource usage, or a minimum level of spend, for one or three years, and Google cuts your rate accordingly. That's the whole mechanism. No upfront payment, no reserved capacity you have to provision yourself. You're simply telling Google "we will use at least this much," and it bills you at a lower rate in return.
The trade-off is straightforward but easy to get wrong. Lock in too rigidly and you pay for capacity you don't use. Stay too flexible and you leave discount money on the table. Getting the balance right means understanding the four commitment structures Google actually offers, because they behave very differently once you're locked in.
Resource-based CUDs are the original, most granular form. You commit to a specific quantity of vCPUs, memory, GPUs, or Local SSD in a specific region, tied to a machine family. These carry the deepest discounts because Google can plan capacity around your exact commitment. The catch is specificity: commit to N2 vCPUs and memory, and that commitment doesn't help you if you migrate half your fleet to C3 or C4 next year.
Compute flexible CUDs solve that rigidity problem. They're spend-based rather than resource-based. Instead of committing to specific vCPU and memory counts on a specific family, you commit to a dollar amount per hour, and Google applies the discount across eligible compute usage regardless of machine family, series, or even region in many cases. The Compute Engine CUD documentation confirms compute flexible CUDs and Flexible Savings Plans both operate on this spend basis, which is precisely why FinOps teams increasingly prefer them for evolving workloads. You lose a little discount depth compared with resource-based commitments, but you gain the ability to shift architecture without stranding your commitment.
Flexible Savings Plans (FSPs) are the newer, more elastic option built for workloads that swing hard within a month. An FSP gives you a monthly entitlement window rather than a fixed hourly commitment spread evenly across a year. According to the FSP documentation, that entitlement resets every month, and whatever you don't use in a given month is gone. There's no rollover. That makes FSPs a poor fit for perfectly steady workloads (a resource-based CUD will beat them on discount depth) but an excellent fit for workloads with seasonal spikes, batch processing bursts, or unpredictable scaling events, where a rigid annual average would either overcommit or undercommit most months.
Service-specific spend-based commitments extend the same logic beyond Compute Engine. GKE Autopilot, Cloud Run, and several other services support their own spend-based commitment structures, following the same core principle: commit to a minimum spend level for a term, receive a proportional discount.
The practical shorthand most FinOps teams settle on:
- Predictable, long-lived workloads on a fixed machine family โ resource-based CUDs for maximum discount depth
- Workloads that will likely change machine family or region within the term โ compute flexible CUDs
- Workloads with sharp monthly variability (marketing spikes, seasonal retail, ML training bursts) โ Flexible Savings Plans
- Non-Compute-Engine services with steady baseline usage โ the relevant service-specific spend-based commitment
Eligible resources and realistic discount rates
Resource-based CUDs cover vCPUs, memory, GPUs, and Local SSD, and this is where the deepest rates live. According to Google's Compute Engine CUD documentation, resource commitments can pay discounts of roughly 55% or higher, with memory-optimised series often sitting at the top of that range because Google prices scarcity into the discount curve.
Statistic Callout: Compute flexible CUD discount rates vary by machine series, with general-purpose series typically offering substantial discounts for one- and three-year terms, and some memory-optimised series offering deeper discounts on longer commitments, according to Google Cloud's pricing documentation.
That roughly 18-point gap between one-year and three-year rates on the same machine series is the single number every FinOps team should sit with before signing anything. A three-year term can significantly increase your discount compared to a one-year term, but it also extends the period during which your architecture needs to remain compatible with the commitment.
Eligibility has some important edges worth knowing before you build a purchase plan:
- Not every machine series supports resource-based commitments. Newer or specialised families sometimes launch with flexible CUD support only, so check current eligibility before assuming a series qualifies.
- GPU commitments follow their own rules and are generally scoped more tightly by region and accelerator type than standard vCPU/memory commitments.
- Local SSD discounts are typically bundled into resource-based commitments rather than sold as a standalone line, so don't expect a separate Local SSD-only commitment product.
- Sole-tenancy nodes carry a premium on top of standard machine pricing, and that premium interacts with CUD discounts differently depending on whether the commitment is resource-based or flexible. Model sole-tenancy costs separately rather than assuming the same discount percentage applies cleanly.
Service coverage beyond Compute Engine has expanded steadily. GKE workloads running on Compute Engine nodes are covered by standard Compute Engine CUDs, while GKE Autopilot and Cloud Run rely on their own spend-based commitment structures rather than resource-based ones, since neither service exposes the underlying VM resource units that resource-based commitments are built around.
How Google applies your commitment against the bill
Google doesn't ask you to nominate which usage a commitment covers. It's automatic, and it follows a fixed order every billing cycle.
- Resource-based CUDs apply first. If you've committed to a specific vCPU and memory quantity in a region and family, Google applies that discount to matching usage before anything else is considered.
- Flexible commitments apply next, covering eligible spend that resource-based commitments didn't already claim. This includes compute flexible CUDs and any active Flexible Savings Plan entitlement for that month.
- On-demand pricing and sustained use discounts (SUDs) apply last, covering whatever usage remains uncovered by any commitment. SUDs are the automatic discounts Google applies for sustained monthly usage regardless of commitments, so they act as a background safety net rather than something you purchase.
Commitment sharing changes how far your coverage reaches. By default, CUD sharing lets a commitment purchased at the billing account level apply across every project under that account, rather than being locked to the project that bought it. This matters enormously for organisations running dozens of projects: a single well-sized commitment purchased centrally can cover matching usage anywhere in the billing hierarchy, not just in the project where finance happened to click "buy."
FSP mechanics deserve a closer look because the monthly window trips people up. An FSP entitlement resets at the start of each calendar month and applies hour by hour as usage occurs. If your entitlement is $500/hour equivalent and you use $600/hour worth of eligible compute in a given hour, $500 gets the discounted rate and $100 spills over to on-demand pricing for that hour. Unused entitlement from a quiet hour doesn't carry forward to cover a busy one later in the same day, and unused entitlement at month end simply disappears.
Pro Tip: Pull your billing export for a single representative week before sizing an FSP. Plot hourly eligible spend rather than daily or monthly totals. Because entitlement applies hour by hour with no rollover, a workload that looks smooth on a daily chart can still be spiky enough hour by hour to waste a chunk of monthly entitlement in quiet periods.
Consider two simplified scenarios on the same $10,000 monthly compute spend. In scenario one, usage is flat at roughly $14/hour throughout the month. A resource-based CUD sized to that baseline covers nearly all of it, at the deepest discount tier available. In scenario two, usage swings between $8/hour overnight and $28/hour during business hours. The same fixed resource-based commitment either overcommits (paying for capacity used only during the day) or undercommits (leaving daytime spend exposed to on-demand rates). An FSP sized to the median hourly spend captures more of that variability without stranding capacity during the quiet hours.
Buying and managing your CUD commitments
You can purchase commitments through the Google Cloud Console, the gcloud command line, or the REST API, and all three end up creating the same underlying commitment resource. Most FinOps teams start in the Console for visibility, then move recurring purchase workflows to gcloud or REST once the process is standardised.
The timing rule that catches people out most often relates to activation, not purchase. According to Google's spend-based CUD documentation, a commitment purchased more than ten minutes before the end of an hour activates at the start of the next hour. Purchase it inside that final ten-minute window, though, and activation slips to the start of the hour after next. That one-hour difference sounds trivial until you're trying to time a purchase against a specific billing cutover or a month-end deadline, and it's worth checking against your account's actual timezone rather than assuming UTC behaviour.
A few practical rules worth building into your purchasing process:
- Buy commitments at the billing account level when you want coverage to extend across multiple projects; buy at the project level only when you specifically need to isolate a commitment to one workload.
- Confirm CUD sharing is enabled at the billing account before assuming a centrally purchased commitment will cover usage in every linked project.
- Choose the term (one year vs three years) based on how confident you are in the underlying architecture staying stable, not purely on which rate looks better on paper.
- Restrict purchase permissions to a small, named group. A commitment is a multi-year financial obligation, not a configuration change, and it shouldn't be purchasable by anyone with general project-editor access.
Governance matters more here than almost anywhere else in cloud cost management, because a bad CUD purchase can't simply be rolled back the way a misconfigured autoscaler can.
Sizing your commitments and calculating expected coverage
Start with a billing export, not a guess. Pull at least ninety days of Cloud Billing export data into BigQuery and isolate the eligible SKUs, meaning the Compute Engine, GKE, or Cloud Run usage that CUDs can actually discount. Ninety days gives you enough history to separate genuine steady-state usage from short-term spikes, seasonal noise, or a one-off migration project that temporarily inflated spend.
- Identify your steady-state floor. Sort your eligible hourly (or daily) spend and find the level that's exceeded roughly 90 to 95% of the time. That floor, not your average or peak, is the safest basis for a resource-based or compute flexible commitment.
- Set a conservative coverage target. Commit to somewhere below that floor, not at it. Market data compiled from GCP pricing benchmarks shows most enterprises achieve moderate coverage of eligible compute spend, while FinOps-mature organisations achieve higher coverage levels. That gap is almost entirely a function of measurement discipline and review cadence, not company size.
- Convert your target into a commitment size. If your steady-state floor is $20/hour in eligible spend and you're targeting 70% coverage, you're sizing a commitment around $14/hour, leaving the remaining $6/hour to on-demand rates and SUDs.
- Model the savings, not just the coverage. At a three-year compute flexible rate of roughly 46%, a $14/hour commitment sustained across a month (roughly 720 hours) saves in the region of $4,600 that month compared with paying on-demand for the same usage, before accounting for SUDs on the uncovered portion.
Statistic Callout: The difference between the typical 55 to 65% coverage most enterprises achieve and the 80 to 90% seen at mature organisations, per industry benchmark data, usually comes down to review cadence rather than workload type. Teams that revisit sizing quarterly close that gap faster than teams that set commitments once and move on.
As a rule of thumb, size resource-based commitments to your genuinely stable, architecture-locked workloads first, then layer compute flexible or FSP commitments on top to catch the remaining eligible spend that's likely to shift shape over the term. Mixing the two deliberately, rather than choosing one exclusively, is usually how mature teams close the coverage gap without overcommitting.
Where commitments go wrong and how to protect against it
The most common way CUD money gets wasted isn't a bad initial purchase. It's drift. A commitment sized correctly on day one against a workload that gets rearchitected, migrated, or decommissioned six months later becomes a stranded cost with no easy exit. Resource-based commitments are especially exposed to this because they're locked to a specific machine family and region.
The FinOps Foundation's own guidance on this is blunt: a governance layer around purchase approval, tagging, and periodic rightsizing is the highest-leverage control available against wasted commitments. That's not a technology problem. It's a process problem, and it's one most organisations underinvest in relative to the size of the financial commitment they're making.
Practical mitigations worth building into your CUD programme:
- Tag every workload covered by a commitment so you can trace whether that workload still exists when the commitment comes up for renewal.
- Assign clear ownership for each commitment, ideally the engineering team whose workload it covers, not just a central FinOps function with no architectural visibility.
- Stagger purchases across quarters rather than committing your entire eligible spend in one purchase, so a single sizing mistake doesn't compound across your whole footprint.
- Reserve Flexible Savings Plans specifically for workloads you already know are spiky, rather than defaulting to them everywhere out of caution.
On contract terms, a few negotiation points matter more than the headline discount rate. Egress pricing, reduction or early-termination rights, and renewal terms all materially change the real economics of a multi-year commitment. Enterprises with substantial GCP spend often negotiate broader Enterprise Agreements that stack with CUDs, and market analysis suggests presenting credible alternative options during that negotiation tends to improve the terms on offer.
Pro Tip: Before renewing any three-year resource-based commitment, check whether the underlying machine series has been superseded. Google periodically ships newer, more efficient series in the same family tier, and renewing into an outdated series locks in savings on hardware you'd otherwise be migrating away from anyway.
Tracking whether your commitments are actually paying off
Four numbers matter more than any others: coverage percentage (eligible spend covered by a commitment), utilisation rate (how much of the committed capacity or spend you're actually consuming), uncovered spend (eligible usage still hitting on-demand rates), and realised savings (the actual dollar difference versus a fully on-demand baseline).
Google's own guidance points to billing export into BigQuery alongside the CUD analysis reports built into the Cloud Console as the two primary data sources for tracking these metrics. The Console reports give you a fast read on current-state coverage and utilisation without any setup. The BigQuery export is where the real analysis happens, because it lets you join commitment data against project, team, or feature-level cost allocation and answer questions the built-in reports can't, such as which specific team's workload is driving an underutilised commitment.
A few habits keep this from becoming a one-off exercise:
- Set an automated alert for utilisation dropping below your target threshold on any active commitment, so drift gets caught in weeks, not at renewal.
- Run a monthly reconciliation between committed capacity and actual eligible usage, flagging any commitment trending toward stranded capacity.
- Report coverage and realised savings to finance in dollar terms, and report utilisation and drift to engineering in workload terms. Each audience needs a different framing of the same underlying data.
- Feed cost allocation data back into architectural decisions, closing the loop between what a commitment is costing and which team's design choices are driving that cost, a link explained in more depth in how cloud total cost of ownership actually gets built.
Reporting cadence matters as much as the metrics themselves. A quarterly review that catches drift early is worth more than a perfectly instrumented dashboard nobody looks at until renewal time.
How Koritsu AI approaches CUD sizing and optimisation
Sizing a commitment correctly means separating genuine steady-state usage from noise, then continuously checking that assumption against reality as architecture changes. That's the exact problem Koritsu AI's platform is built to solve. Our AI agent, Kori, continuously analyses billing data to surface where eligible spend is going uncovered, where a resource-based commitment has drifted out of alignment with current architecture, and where a flexible commitment would fit better than a rigid one.
That analysis feeds into what is called a Savings Opportunity Report: a practical breakdown covering rightsizing opportunities, uncovered steady-state spend that's a candidate for a new commitment, and specific purchase recommendations sized against your actual usage pattern rather than a rule-of-thumb average. Engagements can start with a free assessment, followed by a share of any realized savings, and may move to an ongoing subscription for continuous coverage rather than a one-off review.
What actually determines whether a CUD strategy works
Most of the CUD advice circulating treats the purchase decision as the hard part. It isn't. The hard part is what happens in month four, when someone migrates a workload, a product gets deprecated, or traffic patterns shift and nobody revisits the commitment that was sized against the old reality.
Three rules of thumb consistently separate teams that get real value from CUDs from teams that end up with stranded commitments. First, never size a resource-based commitment against usage you haven't watched for at least one full billing cycle that includes a normal variance event, a deploy, a traffic spike, whatever your version of "normal chaos" looks like. Second, if more than a fifth of your eligible spend sits on a machine family your engineering roadmap plans to retire within the commitment term, that's your signal to use a flexible commitment instead, full stop. Third, treat any commitment purchase as a decision with a renewal date attached, not a one-time transaction. Pausing to run a fresh coverage analysis before every renewal costs an afternoon and prevents most of the drift that wastes committed budget.
If you're ready to act on this, the fastest useful step is mechanical: export the last ninety days of billing data, isolate eligible SKUs, and calculate your current coverage percentage. That single number tells you more about whether your CUD strategy is working than any amount of reading about discount rates.
Get an outside read on your commitment strategy
Sizing a commitment against ninety days of billing data is straightforward in principle. Doing it properly, across every project, every machine family, and every service-specific commitment type, while accounting for what your engineering roadmap is about to change, takes more hours than most FinOps teams have spare in a given quarter. That's the gap a Savings Opportunity Report is built to close.
Koritsu AI runs the assessment for free before anything is charged. Kori, our AI agent, analyses your actual billing history to flag uncovered steady-state spend, commitments that have drifted out of alignment with current architecture, and specific purchase recommendations sized against your real usage pattern rather than a generic benchmark. It's a useful complement to in-house FinOps work rather than a replacement for it. You keep the purchasing decision; we hand you the analysis that makes that decision safer. If a step beyond CUDs turns up too, the same review often surfaces rightsizing and architectural fixes that go further than purchasing discounts alone ever can. We only take a share of the savings we help you realise, verified against your actual billing.
Request your Savings Opportunity Report and get a concrete coverage number to work from before your next commitment renewal.
Sources
Google's CUD overview, Compute Engine CUD documentation, and Flexible Savings Plans docs verify the discount rates and mechanics above. The FinOps Foundation's purchasing guidance informs the governance recommendations, and comparative platform economics are covered in this guide to Microsoft Fabric costs.
- Committed use discounts | Get started
- Purchasing Commitment Discounts in GCP โ FinOps WG
- GCP pricing benchmarks and CUD market analysis