This website uses cookies

Read our Privacy policy and Terms of use for more information.

Presented by:

Want to appear here? Talk with us

Free live session on Sep 23: Finding Cloud Waste That Billing Data Misses

Wasted cloud spend rose to 29% this year, which was the first increase in five years (Flexera).

FinOps Weekly's Victor Garcia gets into where the remaining waste hides inside resources, why billing data never flags it, and how to decide what's safe to delete.

Live only at the end: he breaks down audience-submitted waste on air. 22 minutes, 11am ET / 5pm CET, hosted by Eon.

AI COST OPTIMIZATION
I burned all my tokens researching how to save tokens

A researcher at Quesma burned his entire Claude Max subscription limit in just 30 minutes running a deep research task.

The fix was not spending more money. It was using the tools he already paid for, smarter.

He combined three subscriptions he already had, Claude, Codex, and Antigravity, and let them share memory so no work was repeated.

He matched cheap models to simple tasks and only used the most expensive model for planning and judging.

He made verification a separate step, so one model never checks its own work.

He moved the expensive deep research step to the end of the process, only running it on facts already checked.

The results speak for themselves. Research that used to burn out in 30 minutes now runs for hours.

A process that once used 111 agents and produced nothing now finishes in 22 minutes with 61 agents.

Your harness and workflow can affect cost as much as the model you pick.

Small changes, like how you cache data or compact context, can double your bill by accident.

Track your actual invoice, not just a token calculator, because the gap between the two can be huge.

Before you buy more AI power, check how much power you are already wasting.

AI PROVIDERS
Anthropic Boosts Cache Visibility, Google Expands Overage Controls, OpenAI Launches GPT-6 Astra

Anthropic

Claude Code adds spend-limit and cache tracking, so teams can see costs and cache use in real time.

Claude Code now explains cache misses, helping teams avoid extra token spend from unneeded recaching.

Claude Fable 5.1 launches with clear pricing, listing $10 per million input tokens and $50 per million output tokens for easier cost planning.

Claude Code improves cache reliability and cost tracking, fixing cache resets and adding account details to usage data.

Google

Gemini Enterprise overage controls now work for all invoiced accounts, giving more teams the power to manage extra usage costs.

Gemini Enterprise adds agent performance dashboards, showing speed and error data to help spot costly or unreliable workloads.

Gemini 3.8 Flash is now generally available, offering a faster, lower-cost model option in more regions.

OpenAI

OpenAI launches GPT-6 Astra, a new top-tier model that teams can weigh against cost and capacity needs when planning workloads.

WEBINAR
FINOPS FOR AI

Join this FinOps for AI webinar to learn how to control AI costs, improve visibility, implement governance, and justify AI investments with confidence.

📅 August 27, 2026
🕚 6:00 PM Spain / 12:00 PM ET

VIDEOS & PODCASTS
How to Reduce AWS Costs in Security using FinOps

Learn how to optimize AWS security costs without compromising protection. AWS expert Ihor breaks down common setup errors, FinOps alignment, and budget-friendly AI security strategies.

AI AGENTS
As LLM agent architectures scale in production, context management has become one of the most…

AI agents that use tools and long conversations can rack up huge cloud bills fast.

Every time an agent calls a tool or reads a big file, it stuffs the results into its context window.

That drives up token costs, slows down responses, and can even confuse the AI, making it perform worse on tasks.

Instead of just shrinking text with basic summaries, teams should build systems that keep the exact details an agent needs to make the same decision, while cutting out the noise.

Three ideas stand out for a FinOps audience.

- First, only keep the pass or fail result and error details from tool outputs, not every line of data.

- Second, store large files as small reference tags instead of full payloads, only pulling full data back when needed.

- Third, keep repeated context blocks unchanged so prompt caching still works, which is often the biggest cost saver of all.

Teams that use this approach report cutting agent execution costs by up to 90 percent while keeping accuracy high.

For any business running AI agents at scale, this is a clear reminder that smart context management is now a real cost control lever, not just a technical detail.

AI INFRASTRUCTURE COSTS
LLM Inference Cost on Kubernetes: How to Cut Your Cost per Token

Running AI models on your own servers can quietly burn through cash if nobody checks the math.

The real cost of running LLMs on Kubernetes comes down to a simple ratio.

GPU cost per hour divided by tokens produced per hour.

A single GPU can cost $3 to $5 an hour, and if it sits at low use, every million tokens costs six times more than it should.

Here are the main fixes the article walks through.

Track cost per token using free tools like vLLM and Prometheus, so teams see real numbers instead of guessing.

Batch more requests together, since GPUs can handle dozens of tasks at once for almost the same cost as one.

Right-size the model. Smaller or compressed models often do the job for a fraction of the price.

Scale based on request queue length, not GPU load, since GPU stats can be misleading.

Use spot GPUs and shared GPU slices to cut waste further.

In one example, the same GPU and model went from $3.40 per million tokens down to $0.25 simply by tuning settings.

The takeaway is simple. Measuring cost per token, then adjusting one lever at a time, is what actually lowers AI infrastructure bills, not fancy setups or more hardware.

AI COST VISIBILITY
Observability for AI Agents: Beyond Langfuse Traces

Ever get a cloud bill spike and have no idea which team or workload caused it? That is exactly what is happening with AI agents right now.

One platform engineer found that a single user session cost $12, but there was no easy way to prove it. His tools could not connect the dots between what the AI did, how healthy the servers were, and what it actually cost.

Some key lessons for cost teams:

A single user asking an AI agent to "analyze everything" caused token use to jump 10 to 15 times normal.

That spike ate into shared usage limits and caused slowdowns for other users.

Cloud billing tools show total AI spend by model, not by user or session.

Teams have to build their own tracking to connect a costly AI response to the actual dollar amount.

A small technical fix, reordering how prompts were built, raised cache efficiency from 70 percent to 85 percent.

That one change created a real, measurable savings across thousands of queries.

You cannot control AI costs if you cannot see what is driving them.

Technical monitoring and cost monitoring are not separate problems anymore.

For any team running AI at scale, this is a clear signal to build cost visibility into AI systems now, before usage grows and the bill becomes a surprise.

🎖️ MENTION OF HONOUR
Why The AI Return Shows Up Only When The CFO Owns It

A new study found only 2% of companies give CFOs clear ownership of AI value.

But those companies were far more likely to see real returns from AI, about 76%, compared to roughly half when a CIO or CTO owned the outcome instead.

Approving a budget is not the same as owning the result. Boards ask finance what AI is doing for the business, not IT.

The fix is treating AI spend like any other big cost.

Every AI project should name one clear financial goal before it gets funded, like lower borrowing costs or faster collections.

If it cannot be tied to real money, it is not ready to launch.

Finance teams that check results often, some every quarter or even monthly, catch problems early.

The hardest part is being willing to cut tools that look good in demos but cost more time to fix than they save.

AI spending will keep growing.

Someone needs to own whether it actually pays off, and that job belongs in finance.

Save 20% on AI Value & FinOps Certifications

The job market is hungry for certified professionals who can prove results. Don't let your company's budget leak because of a lack of specialization.

Use code: FINOPSWEEKLY_20 to get an instant 20% discount on the most prestigious certification bundles:

  • FinOps for AI

  • FinOps Certified Practitioner

  • FinOps Certified Engineer

  • FinOps Certified FOCUS Analyst

Liked the Newsletter?

Share your thoughts!

Login or Subscribe to participate