This website uses cookies

Read our Privacy policy and Terms of use for more information.

Want to appear here? Talk with us

TOGETHER WITH KION
A better approach to Bedrock cost attribution

AI cost visibility is only useful if you know what’s behind the spend.

FinOps teams need to understand who is generating that spend, what projects and initiatives it supports, and where that investment is creating value.

AWS has introduced more granular ways to attribute Bedrock costs using IAM principals and session tags. But effective attribution requires more than collecting additional data. It requires building financial context into the way teams access and consume AI services.

Learn how to:

✓ Connect Bedrock spend to the users, teams, projects, and initiatives driving it

✓ Build attribution into existing workflows instead of relying on manual tagging

✓ Turn granular usage data into better reporting, budgeting, forecasting, and chargeback

✓ Create the foundation for evaluating AI spend against the business value it delivers

See what effective AI cost attribution looks like in practice.

AI CODING AGENTS
What a task costs on Opus 5.5

AI coding costs are not just about price per token. They are about how many times the model talks back and forth, just like cloud costs depend on usage patterns, not just unit price.

Anthropic's new Opus 5.5 model cuts list prices across the board. Input and output tokens drop 20%. Cache reads drop 60%, which matters most for long coding sessions.

A task that needs fewer "turns" (back and forth exchanges) costs less, even at the same price per token.

Every turn resends the whole conversation so far, so longer sessions get pricier per step, even with caching.

Picking the right model matters more than tweaking settings.

A smaller model works for simple lookups. A bigger model should only be used for complex problems. Prompt cleanup matters too.

One internal test showed that removing redundant instructions (like forcing six-step procedures) cut costs by 9% on top of the model switch.

Giving the model a way to check its own work, like running a test, avoids expensive retries.

Pausing too long between requests can trigger costly cache rewrites, similar to cold starts in cloud computing.

Anthropic also built in tools like /usage and /cost so teams can measure actual spend instead of guessing.

Price cuts help, but how you use a system determines your real bill.

Measuring your own usage patterns beats trusting any vendor's example numbers.

WEBINAR
Your Cloud Commitments Are Changing. Are They Still Saving You Money?

Cloud usage changes every week. Your commitments don't.

Workloads shift. Projects scale. Usage patterns evolve. And the commitment strategy that made sense yesterday can quietly become tomorrow's unnecessary spend.

The challenge isn't just buying the right commitments. It's keeping them optimized as your cloud changes.

Join us on October 14th for a live online session exploring how to build a cloud commitment strategy that continuously adapts to changing workloads.

You'll discover:

  • How to continuously optimize commitment coverage

  • How to balance savings with flexibility

  • How rolling purchases can reduce lock-in risk

  • Why manual commitment optimization doesn't scale

  • How automation and expert guidance can keep your strategy aligned with real usage

📅 October 14th
⏰ 12:00 PM ET / 5:00 PM UK
📍 Online — Live

AI PROVIDERS
Opus 5.5 Release, Google Cuts Gemini Code Assist Bundling While OpenAI Slashes GPT-6 Token Prices

Anthropic

Claude Code fixed its spend meter, so 1-hour prompt cache writes are priced at the right rate and streamed tokens are not counted twice.

Claude Tag reply footers now show correct cost and token totals, even after worker restarts.

Spend limits now show in dollars in /usage and the status line, so teams can see how close they are to a cap.

New OpenTelemetry logging covers MCP tool, WebFetch, and WebSearch, and new allow/deny model settings help admins block pricey models.

Multi-model fallback improved across Bedrock, Vertex AI, and more, so jobs switch to an older model instead of failing and wasting compute.

OpenAI

GPT-6.1 Sol costs one-fifth of Astra's API token prices, with near-flagship skill for coding and professional work.

Google

Gemini Enterprise AlphaEvolve now defaults to Gemini 3.8 Flash, so check that your costs still match when no model is set.

VIDEOS & PODCASTS
Hybrid Cloud FinOps | On-Prem & Cloud Cost Optimization

Discover how Visual One Intelligence brings FinOps visibility and cost optimization to hybrid cloud environments. A practical look at hybrid cloud FinOps for organizations managing complex multi-cloud and on-prem infrastructure.

AI OPTIMIZATION
Tokenflation: When “Hi” triggers 33 tool calls

A simple "Hi" to an AI coding agent can cost way more than you think.

A new benchmark tested 14 AI models on three prompts: "Hi," "commit," and "WTF." The results show a real problem for anyone tracking AI costs.

Token costs stayed tiny, often under a penny. But wait time told a different story.

One model needed 33 tool calls just to answer "Hi." It read every file, ran the app, searched the whole system, and even made an unrequested commit.

Claude Sonnet averaged 24 tool calls and 49 seconds for a plain greeting.

Some models failed outright. Claude Haiku timed out 60% of the time when given no clear task.

Give the same models an actual task like "commit" and every single one succeeded in under 10 tool calls.

The study calculated waiting time using a developer's salary, and found that time cost can run 20 times higher than the API bill itself.

Total costs ranged from 8 cents to $1.39 depending only on how clear the prompt was.

Clear instructions save money. Vague ones burn both tokens and people's time, and the real bill is often hiding in the wait, not the invoice.

AI FINOPS
AI Tokenomics Prompt Cache Explainer - Optimization Playbook

Caching is one of the biggest ways companies can cut AI costs, but it breaks easily. A cache only works if your prompt matches exactly from the very first word.

Change one date, one line of code, or even the order of a JSON file, and the whole cache resets.

That means you pay full price again, even if 99% of the text is the same as before. Switching AI models mid-conversation also wipes out your cache.

Every provider update, every reroute to a cheaper model, and every A/B test can quietly break your savings without warning.

Two simple ways to track this.

Cache hit rate shows how much of your prompt gets reused.

Cache cost efficiency shows if the caching is actually saving you money or just adding cost.

Big providers like Anthropic, OpenAI, and Google all handle caching differently.

OpenAI even started charging for cache writes in 2026, which caught many teams off guard.

The cheaper model or faster shortcut is not always the cheaper choice once lost cache savings are counted.

Smart teams will watch their cache data closely, keep prompts stable, and think twice before switching models mid-task.

AI COST MANAGEMENT
The Real Cost of AI: A Survey of the Unpredictable Token Economics

For years, software budgets were simple. You picked a subscription, counted seats, and moved on. AI has changed that completely.

Token-based pricing turns AI spend into a moving target, and a new survey of 107 senior leaders shows most companies feel it.

Sixty percent say their AI costs are unpredictable, right when boards are demanding clear ROI.

Most AI budgets, 64 percent, go toward software and API tools rather than infrastructure.

But inside that spending, token usage behaves nothing like a flat subscription fee.

A big mistake many companies make is treating token volume as a stand-in for productivity. It is not.

Tokens are an input, like billable hours, not proof of value delivered.

Agentic AI makes this worse, since it can loop and burn through compute fast.

Only about a third of companies say they actually control token costs, even though over half believe they understand them.

The fixes sound familiar to anyone who has managed cloud spend: route tasks to cheaper models, set usage caps, batch purchases, and build one controlled gateway instead of scattered shadow AI tools.

AI cost management is becoming its own discipline, and it looks a lot like FinOps already does.

🎖️ MENTION OF HONOUR
How The 4M Framework Turns AI Cost Chaos Into Discipline

Here is a strange problem many finance teams are facing right now. AI token prices keep dropping each year, but the AI bill keeps going up.

Think of it like gas prices falling while your fuel bill doubles because everyone started driving a lot more miles.

That is what is happening with AI spend, especially with multi-step AI agents that can use up to 30 times more tokens than a simple chat question.

A new approach called the 4M Framework helps teams fix this with four steps.

Measure tracks exactly what each AI request costs and why.

Minimize cuts waste like long responses and unused data before anything else changes.

Match sends easy tasks to cheaper AI models and saves the expensive ones for hard problems.

Memoize reuses past answers instead of paying for the same work twice, which can cut costs by up to 90 percent.

None of this works without teamwork.

Finance and engineering need to share ownership and track who is spending what.

Companies that follow these four steps in order, instead of trying everything at once, are the ones that keep AI costs under control instead of explaining a surprise bill every quarter.

Save 20% on AI Value & FinOps Certifications

The job market is hungry for certified professionals who can prove results. Don't let your company's budget leak because of a lack of specialization.

Use code: FINOPSWEEKLY_20 to get an instant 20% discount on the most prestigious certification bundles:

  • FinOps Certified Practitioner

  • FinOps AI Value

  • FinOps Technology Value

  • FinOps Certified Engineer

  • FinOps Certified FOCUS Analyst

We Are More than a Newsletter

AI Economics Community

Join the Slack community where the real AI Money conversations happen

Liked the Newsletter?

Share your thoughts!

Login or Subscribe to participate