This website uses cookies

Read our Privacy policy and Terms of use for more information.

Presented by:

Want to appear here? Talk with us

Free live session on Sep 23: Finding Cloud Waste That Billing Data Misses

Wasted cloud spend rose to 29% this year, which was the first increase in five years (Flexera).

FinOps Weekly's Victor Garcia gets into where the remaining waste hides inside resources, why billing data never flags it, and how to decide what's safe to delete.

Live only at the end: he breaks down audience-submitted waste on air. 22 minutes, 11am ET / 5pm CET, hosted by Eon.

AI COST MANAGEMENT
The Ultimate Guide to AI Token Costs

AI tools are quickly becoming a new line item on the cloud bill. Teams are running agents like Claude Code and Codex around the clock.

Without guardrails, token costs can spiral fast.

This GitHub project is a curated list of tools built just for that problem.

It groups resources into what FinOps teams already know: monitoring, optimizing, and governing spend.

For monitoring, there are dashboards like ccusage and Claude Code Usage Monitor that track token use and cost in real time.

Bigger platforms like Datadog and Grafana Cloud now offer AI cost observability too.

For cutting costs, the list covers caching tools that avoid paying for repeat questions.

It also covers model routing, where easy tasks go to cheap models and hard tasks go to expensive ones.

One example claims 90 percent of the quality at just 4 percent of the cost.

For governance, gateways like LiteLLM and Portkey let teams set hard dollar limits on AI spend.

Some tools even auto-block requests once a budget cap is hit.

This list is a practical starting point for any team trying to keep AI costs under control before they become a budget problem.

AI PROVIDERS
Anthropic syncs Claude Code pricing, Google expands Gemini billing access, OpenAI launches GPT-6 Astra and Agents API

Anthropic

Claude Code now syncs gateway pricing with spend data, so cost views match what the gateway actually charges. This makes spend tracking easier for teams using app gateways.

A new max effort setting caps how much reasoning Claude Code uses, giving admins better control over inference costs. Prompt-cache fixes also cut down on repeated data transmission.

Claude Code now explains why prompt-cache misses happen, showing causes like tool or system prompt changes. This helps teams find and fix costly cache misses faster.

OpenAI

GPT-6 Astra launches with better cost controls, including adjustable reasoning effort and cached prompt prefixes. These features help manage token use in long, complex tasks.

The new Agents API manages long-running agent workloads automatically, cutting the work needed to run and coordinate agent infrastructure. This can lower operational overhead for teams building agents.

Google

Gemini Enterprise's pay-as-you-go billing is now open to all invoiced accounts, removing earlier access limits. More teams can now use flexible, usage-based billing.

WEBINAR
FINOPS FOR AI

Join this FinOps for AI webinar to learn how to move beyond simply tracking AI spend and start managing AI Economics, balancing cost, value, and scale.

Discover how to identify AI waste, connect AI costs to business outcomes, measure ROI, and control the full range of AI cost drivers, from tokens and inference to compute, network, and data.

📅 September 17, 2026
🕚 6:00 PM Spain / 12:00 PM ET

VIDEOS & PODCASTS
Why your FinOps Strategy is Failing

We explore Breaking Silos in FinOps and how organizations can bring Engineering, Product, Finance, FinOps and Leadership together to improve cloud cost management.

AI ROI
Calculating the Return on Investment (ROI) of AI

Figuring out if your AI spending is actually paying off can feel like guesswork.

This AWS post breaks down a simple way to calculate real ROI on AI investments.

Here is the core idea:

  • First, sort your AI use into two buckets: external (tied to revenue) and internal (tied to productivity, like coding assistants).

  • Next, add up the full cost, not just the AI tool itself. Include storage, data transfer, monitoring, and agentic AI costs.

  • Use tools like Amazon Bedrock's tagging and IAM tracking to see exactly which team or app is spending what.

  • Pick one clear business metric per use case, like bugs fixed or deployment speed, and make sure it cannot be gamed.

  • Divide your AI cost by that outcome to get a Cost per Outcome. This is your building block for ROI.

Think of it like tracking cost per repair in a shop. If your AI cost per bug fixed drops over time, you are getting more value for the same spend.

Know your cost per outcome, tie it to real business value, and you turn AI from a mystery expense into a measurable investment.

GPU COST OPTIMIZATION
Azure GPU Inference at Scale: A Cost Optimisation Framework by Workload Type

Most companies think GPU bills are a pricing problem. The real data says otherwise.

Across more than 23,000 production Kubernetes clusters, average GPU use was just 5 percent. On Azure Kubernetes Service, it was only 2 percent.

Buying discounts like Reserved Instances or Spot capacity does not fix this. It just makes the waste cheaper.

GPU inference is not one workload, it is three: batch, real-time, and streaming. Each needs different scaling signals and purchasing plans.

Batch jobs can use Spot capacity and scale on queue depth.

Real-time services need a warm baseline, since starting from zero can take minutes.

These should scale on request latency, not raw GPU use.

Streaming workloads should scale on event backlog, since GPU usage alone hides how fast work is piling up.

Tools like MIG, KEDA, and Spot each solve a different problem.

MIG fixes idle GPUs. Spot lowers cost for flexible work. KEDA matches capacity to real demand.

None of these tools replace good workload design.

Classify the workload first. Fix utilization next. Only then choose how to buy the capacity.

A discount on a poorly designed system just locks in the waste at a lower price.

AI AGENT COST OPTIMIZATION
When AI Agents Work, But Still Burn Your Budget

AI agents that work fine can still cost you more than they should. A new AWS guide looks at what happens after an AI agent is up and running.

The problem is not crashes or errors. It is slow answers and memory that keeps growing until sessions fail. These issues are quiet.

They do not trigger alarms, but they slowly drain user trust and your budget. Here is what stood out for cost teams.

Token usage drives cost directly.

Longer AI responses take more time and cost more money, since models generate one token at a time.

Sequential processing wastes both time and spend.

Running tool calls one after another instead of at the same time can double your latency and cost for no good reason.

Long chat sessions can quietly build up memory.

Without limits, this leads to failed sessions and wasted compute.

The fix is proactive monitoring.

AWS tools like CloudWatch and AgentCore Observability help teams catch slowdowns and memory bloat before they turn into real cost problems.

AI agents need cost guardrails just like any other cloud resource, built in from day one, not bolted on after the bill arrives.

🎖️ MENTION OF HONOUR
AI Cost Management: FinOps Lessons Before Your Renewal

Rolling out an AI platform feels like a win when people actually use it. But that win can turn into a budget shock fast.

AI tools are sold like normal software, with seats and a yearly price. But under the hood, they run like cloud services.

Costs go up based on how much people actually use them.

That means a fixed budget can meet a variable bill, and the gap shows up around month seven, right when adoption looks great.

The piece offers real fixes. Model usage on a curve, not a straight line. Expect a small group of users to drive most of the consumption.

Treat add-ons like document or spreadsheet integrations as their own cost line, not a free feature.

Decide what happens when you hit your usage cap before you actually hit it.

- Get real visibility into who is using what, and how much.

- Track cost per active user over time, not just the total spend.

- Start renewal talks ninety days out, not thirty.

Heavy usage is not the problem, it is the reason the tool is worth paying for.

The real fix is watching the meter early, so you can fund success instead of stopping it.

Save 20% on AI Value & FinOps Certifications

The job market is hungry for certified professionals who can prove results. Don't let your company's budget leak because of a lack of specialization.

Use code: FINOPSWEEKLY_20 to get an instant 20% discount on the most prestigious certification bundles:

  • FinOps for AI

  • FinOps Certified Practitioner

  • FinOps Certified Engineer

  • FinOps Certified FOCUS Analyst

Liked the Newsletter?

Share your thoughts!

Login or Subscribe to participate