Cloud Cost Optimization: Why the Market Split When Waste Came Back

Amazon Web Services, Microsoft Azure, and Google Cloud already sit on the authoritative bill. AWS Cost Explorer and the Cost and Usage Report, Azure Cost Management and Billing, and Google Cloud Billing are where every serious practice starts, because they are free enough, complete enough for that provider, and close to the identity and resource model.
nOps automates scheduling, rightsizing, and spot use, with a commercial model that can be paid from savings. Spot by NetApp acts on placement and capacity. PointFive is in the business of finding waste that can be closed, not only displayed. AWS Compute Optimizer remains the native baseline that many automation tools still compete with or ingest.

What changed

Cloud cost optimization used to mean a monthly export, a tagging campaign, and a reserved-instance purchase. That was a reasonable response when most of the money sat in a handful of virtual machines on one provider, owned by a central infrastructure team.
Cloud cost optimization is not one category. It is three jobs that get sold under one phrase.
Because these tools live on the resource path, they cannot report spend they are not connected to. SaaS, many data platforms, and most model-provider bills are outside the loop. Automation without allocation also produces the wrong wins: it will shrink a sandbox and leave the unowned production GPU fleet untouched, because the sandbox was easier.
What good looks like is not a pretty namespace pie chart. It is allocation that matches how the platform is actually billed (including idle resource, attached disks, ingress, and GPU time), with costs pushed back to the same identity model the rest of FinOps uses. If cluster cost lives in a tool that finance never opens, you have two truths.
The map below is the market as a buyer actually meets it. The sections after it walk through each row, including what that row cannot do.
If most of the money is still in one hyperscaler and one team owns it, native tools plus commitment hygiene are the whole first year. Buying an independent platform on day one, before tags and account structure exist, produces a prettier version of confusion.

The problem is not the invoice

What good looks like here is boring: CUR or equivalent exports landing in a warehouse, consistent tags or labels on new resources, budgets and anomaly alerts that page a person, not a shared inbox. Native tools fail when the company pretends they are a multi-cloud financial system. They also fail on Kubernetes, where the billable object is a node or a disk and the responsible team is a namespace. They do not see SaaS, and they do not see AI tokens purchased from a model vendor.
The first job is rate: pay less for the same usage, through commitments, spots, and enterprise agreements. The second is usage: run less, or run smaller, through rightsizing, scheduling, idle cleanup, and workload placement. The third is attribution: know what a customer, team, feature, or inference actually costs, so someone with a budget will care.
Kubecost, now part of IBM, became the default commercial answer for many platform teams. OpenCost, the CNCF project, is the open-source allocation model those conversations now share. Finout and nOps extend into cluster cost because their buyers refused to keep Kubernetes in a side spreadsheet.
The structure here is not a stack. Later layers do not sit on earlier ones the way a control plane sits on agents. The structure is a set of competing levers, with a few real dependencies hiding inside them.

The market thesis

Watch the seams, not the logos. IBM now holds Cloudability and Kubecost, which is a coherent answer to “finance plus cluster” and still not an answer to unit cost or to SaaS AI. Flexera has bought its way across two more meters in a single January, taking commitment automation with ProsperOps and lakehouse cost with Chaos Genius; acquisition moves faster than a shared data model, and buyers should ask which is in place. Broadcom’s handling of Tanzu CloudHealth will matter to a large installed base that did not choose Broadcom. Independents will keep adding automation because reporting-only businesses get squeezed between free native tools and platforms that can act.
These platforms exist because a million cloud bill is also a control problem. Someone has to produce a showback that business units accept, a forecast that planning will use, and a policy that new accounts cannot evade. Flexera reports 63% of organizations now have a FinOps team; the enterprise suite is what those teams buy when the audience is the CIO and the CFO, not a Slack channel of SREs.
If the painful number is Kubernetes, start in the cluster. Enterprise FinOps platforms that cannot ingest in-cluster metrics will allocate by node and then argue with platform engineering for the rest of the contract.
The recommendation siblings matter as much as the charts. AWS Compute Optimizer, Azure Advisor, and Google’s recommender APIs propose rightsizing and idle cleanup from inside the same account. Commitment products (Savings Plans, Reserved Instances, Azure reservations, Committed Use Discounts) are still the largest single rate lever most companies have.

How the market is taking shape

That combination is the story. The discipline got more serious. The bill got harder to hold still.
This layer has a clean success metric: effective savings rate on eligible compute. It also has a clean failure mode. If usage is the problem (idle, oversized, forgotten), a better rate makes the waste cheaper and more durable. If product-level attribution is the problem, a Savings Plan will not tell a general manager whether a feature should exist.
The invoice is late, coarse, and politically useless on its own.
This is not a collapse of FinOps. It is complexity growing faster than the controls built for the last generation of bills.

  • Inside the provider, where the data is complete for one cloud and incomplete for everything else.
  • Above the bill, where independent platforms normalize invoices, attach business context, and argue about unit cost.
  • On the commitment, where specialists buy, sell, and float reservations so humans do not have to.
  • In the cluster, where the shared node is the real cost object.
  • On the workload, where the fix is to change running infrastructure, not to interpret a CSV.

Around those constraints, the market sorted itself by where it sits relative to the bill:
Also watch whether “FinOps platform” quietly becomes “technology spend platform.” The Foundation’s own scope now includes SaaS, licensing, private cloud, and data center. That expansion is real in the survey data. It is not yet real in most products’ data models. Buyers who take the mission statement as a feature list will over-buy.

How cloud cost optimization approaches relate
Market Layer What It Solves Typical Buyer Representative Vendors Where It Fits
Native billing and recommendations One provider’s invoice, basic allocation, and first-party rightsizing or commitment advice. Cloud operations or FinOps on a primary hyperscaler. AWS Cost Explorer, Azure Cost Management, Google Cloud Billing Accurate and cheap for one cloud. Breaks on multi-cloud, SaaS AI, and shared clusters.
Enterprise FinOps platforms Chargeback, governance, hybrid and license context, reporting that finance will accept. FinOps, ITFM, and CIO-aligned cost centers in large enterprises. IBM Cloudability, Flexera, VMware Tanzu CloudHealth (Broadcom) Strong on process and allocation. Often slower to act on live infrastructure.
Independent cost intelligence Normalized multi-cloud spend, unit economics, engineering-facing cost of a product or customer. FinOps paired with engineering or product finance. CloudZero, Vantage, Finout, nOps, PointFive Sees across providers. Cannot change a resource it does not control.
Rate and commitment automation Reserved Instances, Savings Plans, CUDs, and spot mix without a spreadsheet owner. FinOps or cloud center of excellence with a large, spiky compute bill. ProsperOps (Flexera), Spot by NetApp, nOps Improves price, not shape. Will not fix idle, oversized, or unowned usage.
Kubernetes cost allocation Splitting a shared cluster by namespace, workload, or team, using in-cluster metrics. Platform engineering and FinOps for container platforms. Kubecost (IBM), OpenCost, Finout, nOps Necessary for cluster economics. Blind to most of the non-container bill.
Usage and workload automation Rightsizing, scheduling, idle cleanup, and placement that changes running resources. Platform or SRE teams willing to let software touch production. nOps, Spot by NetApp, PointFive, AWS Compute Optimizer This is control, not reporting. Scope is only the resources the tool is allowed to touch.

Native tools still win the first cloud

If the painful number is on-demand markup, rate automation pays for itself. Do not let that success substitute for usage work. Cheap idle is still idle.
Estimated wasted cloud spend rose in 2026 for the first time in five years. Flexera’s 2026 State of the Cloud Report, based on 753 cloud decision-makers, puts that waste at 29% of spend, even as more organizations staff FinOps teams and talk about business value rather than discounts.
So the job title stayed “cloud cost optimization.” The object of the job did not.
Two things broke that model at the same time.

Enterprise platforms grew up around finance, not kubectl

The FinOps Foundation’s 2026 survey puts the frontier elsewhere. FinOps for AI is now the top forward-looking priority for the discipline, and AI cost management is the single skillset teams most want to add. Rate work is necessary and increasingly automated. It is not where the practice is straining.
Second, FinOps stopped being only about public cloud. The FinOps Foundation’s State of FinOps 2026 survey, covering 1,192 respondents and more than billion in annual cloud spend, found 98% of organizations now manage some form of AI spend, up from 31% two years earlier. The Foundation’s 2026 mission update puts adjacent scope in plain numbers: 90% manage SaaS, 64% manage software licensing, 57% manage private cloud, 48% manage data center spend. A slice of practices even pull labor into the same system.
The shared origin shows up in the gaps. They are strong on process and weak on the last mile of change. A recommendation to downsize a node pool is not the same as doing it. Kubernetes depth often arrives through a second product (IBM also owns Kubecost) rather than through the financial console. Multi-cloud display is common; identical allocation logic across AWS, Azure, Google Cloud, and a data warehouse is not. Buyers who confuse “we have a FinOps platform” with “we optimize workloads” discover the difference in the second year, when the dashboard is mature and the waste number has not moved.

Independent intelligence vendors made cost an engineering object

Ownership is the other dependency. A recommendation with no named owner is a report. Automation without a change window and a rollback path is an outage with a savings chart attached.
The boundary is control. Seeing a runaway GPU job is not stopping it. Several of these vendors now attach automation or partner with it, which is why nOps appears in more than one row of the matrix. The remaining risk is coverage theater. A true multi-cloud platform allocates the same way on every provider and can join AI platform spend into the same model. A tool that merely paints three invoices in one UI has not done that work. Ask which of those you are buying.
If AI is the new line item, ask where the tokens live. Cloud compute for training and inference can sit in the same FinOps model you already have. Embedded SaaS AI and direct model-provider bills often cannot. The FinOps Foundation has already recorded that 98% of surveyed practices are being asked to manage AI spend; most of the tooling in this article was built for IaaS.
The underlying problem is mismatch. Money leaves through unused capacity, oversized resources, uncommitted on-demand rates, untagged shared services, and new consumption types that the old allocation model cannot name. The people who can change those things (engineers, platform teams, procurement, product) do not share a system of record. Optimization, as a market, is the attempt to close that mismatch.

Rate automation is a different product pretending to be the same conversation

Allocation quality is the dependency that actually bites. Rightsizing advice on untagged, shared, or multi-tenant resources is theater. Commitment purchases made against a blended view of production and sandbox will lock in the wrong shape of spend. Kubernetes tools that cannot see the cluster will allocate by guess. Gateways and request-path AI cost tools will not see spend that never passes through them.
If you treat “cost optimization” as a feature checklist, you will buy a dashboard. If you treat it as three jobs, you can see why the market fractured.
ProsperOps built a business on floating and managing AWS commitments so a human does not babysit coverage versus lock-in. Flexera acquired it in January 2026, alongside Chaos Genius for Snowflake and Databricks cost, and says ProsperOps will continue operating under its own brand. Spot by NetApp (the former Spot.io, inside NetApp) has long sold a mix of spot capacity, elastigroup-style placement, and related rate tricks. nOps includes commitment handling in an AWS-heavy automation suite. The hyperscalers, of course, sell the raw instruments.
First, the unit of spend stopped being the instance. Kubernetes shared a node across teams. Serverless hid the machine. Data platforms billed on consumption that did not map to a cost center. Generative AI moved from experiment to everyday service. Flexera found GenAI at 58% adoption as a public cloud service in 2026, up from 50%, and ranked it the third most widely used public cloud service in that survey. The same report still lists managing cloud spend as a top challenge for 85% of respondents, with 63% now running a FinOps team and 64% claiming they can show value delivered to business units.

Kubernetes broke the bill, so a separate market formed inside the cluster

Shared nodes made classic tag-based allocation a fiction. The honest cost object is the pod, the namespace, the persistent volume, and the load balancer the chart forgot to delete.
This is the only layer that can move the 29% waste number without a human clicking through a ticket for every instance. It is also the layer that creates outages if ownership is fuzzy. The buying requirement is not “AI-powered recommendations.” It is a defined blast radius, a record of what changed, and an identity for the team that will get paged.
Commitment management is where a lot of “we saved 30%” stories actually come from, and it is a poor proxy for the rest of FinOps.
A single-cloud shop with a strong tagging culture can get surprisingly far on this layer alone. That buyer still exists. The market’s noise is generated by everyone else.

Usage automation is where reporting stops being the product

A second group formed around a different complaint: the bill is not how engineers think.
The constraint that will sort this market is authority. If waste is rising while teams and tools multiply, the missing piece is not another chart. It is permission, in production, to change the thing that costs money.
Tools are good at one of these and noisy at the others. That is why there is no obvious single product, and why consolidation in the vendor base has not produced a single buyer motion. IBM can own both a FinOps platform and a Kubernetes cost product and still leave a team short on unit economics. A hyperscaler console can be perfectly accurate for its own bill and still be the wrong system of record for a company that runs two clouds and a data platform.
A line that says “EC2” or “Compute Engine” does not tell a product owner whether a feature is expensive. It does not tell an engineer that a Kubernetes Service of type LoadBalancer has been sitting idle in a forgotten namespace. It does not tell finance whether a Savings Plan is covering the right family of usage, or just producing a comforting coverage percentage. And it does not see a token bill inside a SaaS AI feature that never appears in a Cost and Usage Report.

What those distinctions mean for a buyer

This layer does not optimize the warehouse, the SaaS suite, or the OpenAI invoice. Observability of a cluster without permission to change it also will not stop a retry loop. Platform teams buy Kubecost and then discover they still need a FinOps system of record, or they buy Cloudability and discover they still cannot explain a noisy neighbor on a node pool. Both discoveries are the market working as designed.
A last group treats the cloud as something you change.
Buyers already know they should “reduce waste.” That sentence has been true for a decade. What they need now is a way to decide which waste they can actually touch, with which data, and who is allowed to touch it.
Because they sit on normalized billing plus whatever telemetry they can ingest, they can answer questions native consoles duck: what does this microservice cost in production, and who owns the increase since Tuesday. That is why they show up in engineering orgs that already rejected a six-month tagging program as the only plan.
Waste returning while FinOps headcount rises is a warning about scope, not about effort. AI training jobs, inference endpoints, and token bills do not behave like a fleet of m5.large instances. They spike, they hide inside other products, and they get funded with a political story that is hard to unwind. Any vendor still selling only “rightsizing EC2” is selling last year’s problem.
Organizationally, the FinOps Foundation’s 2026 mission update reports that 78% of FinOps practices now sit inside the CTO or CIO organization, and that practitioners with VP or C-suite engagement show markedly more influence over technology selection than those with only director-level access: 53% versus 24% on cloud service selection, 47% versus 16% on provider selection. That is not a staffing footnote. Cost optimization that reports only to procurement will buy rate products. Cost optimization that sits with engineering can buy usage change. The reporting line predicts the stack.

  • Buying reporting before ownership. Untagged shared accounts make every platform look insightful and none of them actionable.
  • Treating visibility as control. Dashboards do not stop a pipeline. If no tool is allowed to change production, the waste number is a KPI, not a program.
  • Using native coverage metrics as the company score. A 90% Savings Plan coverage rate can coexist with a 29% waste rate. They measure different jobs.
  • Expecting a Kubernetes tool to explain the rest of the bill, or a FinOps suite to explain a node pool. Adjacent is not included.
  • Evaluating vendors on the number of clouds in a screenshot. The test is whether allocation, anomaly detection, and (if promised) automation work the same way on each source, including the AI and data platforms you actually run.

If spend is multi-cloud, or if engineering will not log into a finance console, independent intelligence is the system of record. It still needs a native export underneath it. Nothing in this market replaces the CUR; it interprets it.

What to pay attention to next

A few traps show up so often they are almost a buying process:
CloudZero, Vantage, Finout, nOps, and PointFive all sell some version of cost as a product metric. The pitch varies. CloudZero has pushed unit economics (cost per customer, per feature) harder than most. Vantage grew as a fast, readable overlay on native billing. Finout has leaned into allocation across clouds and Kubernetes. nOps mixes visibility with AWS-centric automation. PointFive sells waste finding with an eye on what can actually be reclaimed.
Those jobs do not replace each other. A company can be excellent at Savings Plans and still hemorrhage money on idle GPUs. It can have perfect showback and still never shut anything off.
The first independent response was built for companies that already had IT financial management. IBM Cloudability (the Apptio product, now inside IBM) is the clearest example: allocation, planning, and chargeback that a finance partner will recognize. Flexera comes from software asset and cloud management and still carries that hybrid and license instinct into FinOps. VMware Tanzu CloudHealth, now a Broadcom product and sold through Arrow Electronics as its global provider, remains the other long-running enterprise console, carried through two renamings since VMware bought CloudHealth in 2018.

Similar Posts