FinOps Inform

Engineers: Three policy as code rules to stop costly cloud waste

Engineering playbook: use policy as code to stop cloud overspend with platform guardrails, CI/CD shift left checks, and FinOps support.

Engineer reviewing cloud deployment policy

Yes: policy as code can prevent and limit cloud spend by enforcing pre-deployment guardrails and shift-left cost checks. The mechanism is straightforward. You write cost rules as code, run them against every plan or pull request, and block or flag anything that breaks the budget before it reaches production. It only works, though, once your tagging and account hierarchy are solid enough to give those rules something accurate to check.


TL;DR:

  • Policy as code should be implemented across multiple layers, including organization-wide controls, pull request checks, and automated budget actions, to maximize effectiveness.
  • Enforcing mandatory tags, resource size limits, and PR-level cost caps can significantly prevent overspending and catch issues early in the deployment process.
  • Rollout of policies should start as advisory, then progress to warning, and finally to blocking, to reduce friction and increase adoption.
  • Proper attribution of savings and regular review of active exceptions are essential to demonstrate policy-driven cost reductions over time.
  • Centralizing cost policies within existing governance frameworks and using provider-agnostic tools help maintain consistency and simplify management across multi-cloud and hybrid environments.

How policy as code works for cost control: core concepts and architecture patterns

Policy as code started as a compliance tool: enforcing encryption, network rules, and access controls automatically instead of relying on manual review. Cost-aware guardrails borrow the same mechanism but check a different question. Instead of "is this secure?", the policy asks "does this cost too much, and does it belong to someone?"

That second question needs inputs that security policies do not. You need current pricing data, a cost estimate generated at plan time, consistent tags, and a clear account or subscription hierarchy that maps resources back to teams and budgets. Without those four things, a cost policy has nothing reliable to evaluate.

There are two places to put the enforcement logic. Platform-native tools such as Azure Policy or AWS Organizations apply rules at the account or subscription level, catching anything that slips through regardless of how it was deployed. Pipeline-level engines such as Open Policy Agent or Sentinel evaluate a Terraform plan before it ever reaches the cloud, giving engineers feedback in minutes rather than after the bill arrives. Cloud architecture guidance advises prioritising platform-native automation and supplementing it with CI/CD pre-deployment validation, which is the layered model most mature teams end up with:

  • Platform-level controls act as an unbypassable backstop across every account.
  • Pipeline-level checks catch problems earlier, before deployment, with faster feedback to the engineer.
  • Both layers read from the same tagging and pricing data, so they never disagree on what something costs.

Where to place policies and how to integrate them

Deciding where a policy lives determines how easy it is to bypass and how fast engineers get feedback. The right answer is usually all three layers working together, each with a distinct job.

  1. Organisation and management account controls. Service Control Policies in AWS Organizations or subscription-level policies in Azure sit above individual teams and cannot be edited by a project engineer. These are your backstop: even if a pipeline check is skipped or a policy engine is misconfigured downstream, an SCP can still block an oversized instance type or an unapproved region.
  2. Shift-left checks in the pull request. Running a cost estimate against a Terraform plan and posting the delta as a PR comment turns an invisible cost decision into a visible one, before merge. Outcomes typically fall into three buckets: ALLOW when the change is within budget, WARN when it needs a second look, and BLOCK when it breaches a hard limit.
  3. Platform budgeting and automated actions. AWS Budgets with Budget Actions, or Azure Policy's remediation tasks, can automatically restrict permissions or halt spend once a threshold is crossed, without waiting for a human to notice the alert.

On the toolchain, the common pattern is a Terraform plan evaluated by Open Policy Agent, Rego, or Sentinel, wired into GitHub Actions or a similar CI system so the check runs automatically on every pull request. AWS Well-Architected guidance recommends a layered approach combining budgets, notifications, IAM controls and Service Control Policies, specifically because a single layer is always easier to route around.

Pro Tip: Start the pipeline check as advisory before wiring it to a hard BLOCK. Engineers trust a system that warns them first far more than one that rejects their first pull request without explanation.

Three concrete policy-as-code examples engineers can adapt

Abstract guardrails are hard to picture. These three patterns cover most of what teams need on day one.

Tag enforcement. Require a small set of mandatory tags (owner, environment, cost centre, project) on every resource at creation time, and reject deployments missing them. Pair this with automated tag inheritance, so a tagged parent resource passes its tags down to child resources automatically, and a remediation pathway that flags or quarantines anything still untagged after a set period rather than deleting it outright.

Instance-size and resource-type limits. Cap the instance families and sizes allowed per environment. Development and staging rarely need anything beyond small or medium instances, so a policy that blocks a memory-optimised, 32-core instance type from being provisioned outside production removes a whole category of accidental overspend.

PR-level cost caps. Estimate the monthly cost delta a change introduces and gate the merge on it.

A change that adds £40 a month gets an automatic ALLOW. One that adds £4,000 gets a BLOCK with a required sign-off from the budget owner.

Common shift-left pattern in Terraform-based CI/CD pipelines

The developer workflow stays simple: open the pull request, the pipeline runs the plan, a bot comments with the estimated delta and a verdict, and the engineer either merges or requests an exception. Sentinel-style policy frameworks can enforce these rules at plan time, covering cost thresholds, instance types, and tags in one evaluation, with enforcement ranging from advisory to hard-mandatory depending on the rule.

  • Tag policies protect cost allocation and reporting accuracy.
  • Instance-size limits protect non-production environments from silent overprovisioning.
  • PR cost caps protect the budget at the exact moment a decision is being made.

Choosing enforcement levels and rolling policies out without blocking velocity

Not every policy deserves the same teeth. Advisory policies log a violation without blocking anything, useful for a new rule you are not yet confident in. Soft-mandatory policies warn loudly and require a documented override, which works well for rules with legitimate exceptions. Hard-mandatory policies block outright and suit only the clearest cases, such as a resource type that should never exist in a given account.

A phased rollout avoids the two failure modes teams hit most often: rolling out too fast and getting bypassed through shadow processes, or rolling out too cautiously and never getting enforcement at all.

  • Start in monitor mode: log violations, change nothing, and see what the data shows.
  • Move to warn: surface violations to the engineer directly, with a clear remediation step attached.
  • Graduate to block only for policies with a low false-positive rate and an approval workflow for genuine exceptions.

Friction kills adoption faster than a strict policy does. Every blocked deployment should come with a specific fix, not just a rejection message.

KPIs and attribution: how to prove policy-driven savings

Policies only earn their keep if you can show what they prevented, not just what they blocked. Track tagging compliance rate, the cost delta of PR merges that were blocked or downgraded, the count of warned versus blocked changes over time, and the reduction in cost anomalies month over month.

Attribution is the harder half. Map the estimated spend a policy prevented against what actually shows up on the bill, and report the comparison on a fixed cadence rather than ad hoc. Cost allocation hygiene, including mapping unassigned resources to owners, produces measurable multi-year reductions in cloud spend and gives policy-as-code programmes a credible baseline to attribute savings against.

Anomaly detection reduced through consistent tagging and allocation is one of the clearest leading signals that a cost-governance programme is working, because it means spend is visible enough to spot problems before they compound.

Practitioner perspective: combining continuous analysis and expert FinOps

Policy as code stops new waste. It does nothing about the waste already sitting in your account from a decision made two years ago by an engineer who has since left. That gap is where continuous analysis and hands-on FinOps expertise earn their place.

Koritsu pairs an AI platform that continuously reviews cloud spend with engineers who investigate why a workload is expensive at the architecture level, not just whether it breached a threshold. Guardrails control what happens next; this kind of review reclaims what already went wrong. Bringing in a specialist tends to make the most sense when chargeback across teams is genuinely difficult, tagging hygiene has decayed over time, or you are managing spend across more than one cloud provider at once.

Integration with existing cloud governance frameworks

Policy as code should not run as a parallel system next to your existing governance. It should plug into the frameworks you already use for security, compliance, and access management, using the same evaluation engine wherever possible.

If you already run Open Policy Agent for security policies, add cost rules to the same policy bundle rather than standing up a second tool. If Azure Policy already enforces naming conventions and network restrictions, cost policies belong in the same policy set, evaluated at the same point in the deployment pipeline. This keeps a single audit trail and avoids a situation where a resource passes a security check and a cost check at two different times, with two different results.

AWS Well-Architected guidance treats budgets, IAM, and Service Control Policies as one layered system rather than separate controls, which is the model worth copying: cost governance is one more dimension of the same policy engine, not a separate discipline with its own tooling and its own review board.

Governance boards that already exist for security and compliance are usually the right place to approve cost policy changes too. Adding a second, cost-specific governance committee tends to slow decisions down without improving them.

Handling multi-cloud and hybrid environments

Multi-cloud makes policy as code harder in one specific way: the same rule has to be expressed differently for each provider's native tooling, since AWS Service Control Policies, Azure Policy, and Google Cloud's Organization Policy Service each use their own syntax and scope model.

The practical fix is to centralise the policy logic in a provider-agnostic engine, Terraform plans evaluated through OPA or Sentinel work across AWS, Azure, and Google Cloud, and treat each platform's native controls as the backstop layer underneath. That way a rule such as "no instance larger than a defined size in non-production" is written once in your policy engine and enforced consistently, while each cloud's own guardrails catch anything that bypasses the pipeline entirely.

Multi-cloud policy enforcement flow

Hybrid environments add a further wrinkle: on-premises infrastructure rarely has the same real-time pricing data available, so cost estimates for hybrid workloads tend to rely on amortised or reserved-capacity figures rather than live pricing. Keep hybrid cost policies simpler and more conservative than their cloud-native equivalents until the underlying cost data improves.

Consistent tagging matters even more here, since it is often the only common thread linking a resource in one cloud to its counterpart in another.

Scaling policies for enterprise-wide adoption

A policy that works for one team's ten repositories does not automatically work for two hundred teams and a thousand repositories. Scaling requires treating the policy set itself as a product, with versioning, ownership, and a change process.

Centralise policy authoring with a platform team, but let individual teams request exceptions or propose new rules through a standard process rather than forking the policy set. Version policies the same way you version application code, so a team can pin to a known policy version while a new rule is being tested elsewhere.

Rolling out gradually by business unit or environment, rather than flipping a switch organisation-wide, keeps the blast radius of a badly tuned rule small. A phased approach, monitor first, then warn, then block, applies at the organisational rollout level exactly as it does for an individual policy.

Managing exceptions and policy drift

Every cost policy will eventually meet a legitimate exception: a genuinely necessary large instance, a short-term burst workload, a migration that temporarily needs more headroom than the rule allows. Handle these through a defined approval workflow rather than letting engineers quietly override the policy in the console.

A cost-governance programme works best when platform teams can grant short-lived elevated permissions for these approved cases, paired with an automated audit that checks the exception has actually expired. Without that expiry check, "temporary" exceptions accumulate silently and the policy set drifts further from what is actually enforced.

Time-boxed policy exception workflow

Policy drift, where the written rule and the enforced reality diverge, is one of the most common reasons governance programmes quietly fail. Schedule a regular review of active exceptions and policy versions, and treat a policy that has been overridden repeatedly as a signal that the rule itself needs revisiting, not just that engineers need reminding.

Cost and security guardrails often share the same enforcement plumbing, which is efficient, but it also means a badly scoped cost policy can create security gaps if you are not careful. A temporary exception granting broader permissions to deploy an oversized resource, for instance, can leave a wider blast radius open if it is not time-boxed and audited.

The safer pattern is to treat any exception mechanism as a security control in its own right: scope it narrowly, log every use, and expire it automatically. The same Service Control Policies and Azure Policy assignments that block cost overruns are frequently the same mechanisms restricting which regions, services, and instance types can be used at all, so a change to a cost policy should go through the same review as a change to a security policy, not a lighter one.

A brief first-person reflection for platform leaders

The real trade-off in policy as code is not cost versus compliance. It is gating versus velocity: every hard block you add slows someone down, so save BLOCK for the rules you are genuinely certain about. Ownership works best shared across platform and finance, not owned by either alone. Start soft, watch the data, and tighten enforcement as your confidence grows.

Koritsu: a practical option for teams needing hands-on help

Guardrails stop new waste, but they cannot find the inefficiencies already baked into your architecture from decisions made months or years ago. That is a different problem, and it usually needs a different kind of review.

Koritsu AI

Koritsu starts with a free assessment and a Savings Opportunity Report that maps where money is actually being lost across your AWS, Azure, or Google Cloud estate. Engagements run on a success fee, a share of the savings actually verified against your bill, with no upfront cost. From there, teams can move onto ongoing Monitor or Advisor plans for continued analysis and hands-on support. A sensible next step alongside any policy rollout is a policy gap analysis paired with a cost-mapping review, so guardrails and cleanup happen together rather than one waiting on the other.

Sources

FAQ

Can you give me an example of policy as code?

A common example is a Terraform plan evaluated by Open Policy Agent or Sentinel before deployment, blocking any resource that breaches a defined monthly cost limit or uses a restricted instance type. Another is a mandatory-tag policy that rejects any resource missing an owner or cost-centre tag at creation time.

Which is the best IaC tool?

There is no single best infrastructure-as-code tool; the right choice depends on your cloud providers, team skills, and existing pipeline. Terraform is the most widely adopted option for multi-cloud policy enforcement through engines like Sentinel or OPA, while Pulumi CrossGuard suits teams that prefer writing policies in a general-purpose language.

What is code as policy?

Code as policy, more commonly called policy as code, means writing governance rules such as cost limits, tagging requirements, or security controls as machine-readable code rather than manual checklists. Those rules are then evaluated automatically at plan time or deployment time, allowing violations to be caught before they reach production.

Is IaC considered DevOps?

Infrastructure as code is a core DevOps practice, since it applies software engineering discipline, version control, review, and automated testing to infrastructure changes. Combining it with policy as code extends that same discipline to cost and compliance rules, keeping governance checks inside the same automated pipeline as the deployment itself.