Put a gateway in front of every model call so each token is tagged to a team and a use case, routed to the cheapest adequate model, capped by budget, and alerted on before the vendor cuts you off.
AI spend stopped being a line item somebody else owned. The State of FinOps 2026 survey found 98 percent of its 1,192 respondents now manage AI spend, up from 31 percent two years earlier — and the billing surface got materially harder in the third quarter of 2026. Microsoft moved its long-running agents (Autopilot, Cowork, Code) and frontier models onto usage-based billing on September 25, 2026, while Copilot Studio sells 25,000-credit packs at $200 a month and disables custom agents when a tenant hits 125 percent of prepaid capacity. Google's Gemini 3.7 and 3.8 Flash input price doubles from $0.75 to $1.50 per million tokens on January 1, 2027, with output going from $3.75 to $7.50, and Gemini 4 Argon's introductory $2/$10 steps up to a standard $4/$20 on a date Google has not yet announced. OpenAI charges GPT-6 Astra requests over 272K input tokens at $20 input and $75 output against $10 and $50 base. Anthropic prices cache reads at $0.25 per million on Fable 5.1 and $0.20 on Opus 5.5 and Sonnet 5.5 against $10, $4 and $2 base input, which makes caching the single largest cost lever for agents — and notes that its current tokenizer produces roughly 30 percent more tokens for the same text than earlier ones, which quietly inflates any per-token comparison. Salesforce bills Agentforce in Flex Credits at $500 per 100,000 with 20 credits per standard action and 30 per voice action. Five vendors, five units, and none of them tag spend to the team that caused it. The agent inventory and governance workflow in this library is about risk; this one is about the bill. The mechanism is a gateway — Portkey (now sold as Prisma AIRS AI Gateway), LiteLLM, or OpenRouter — through which every call passes, so that tagging, routing, budgets, and alerts live in one place you control rather than in five vendor consoles you check after the invoice arrives.
Direct API keys to the Claude API, the OpenAI API, and Gemini; cloud-hosted inference through Amazon Bedrock and Azure OpenAI Service; platform credits in Microsoft Copilot Studio and Agentforce; and SaaS products with usage-based AI add-ons. Pull the last three invoices for each. The first finding is almost always that a majority of spend sits on shared keys with no owner, which is the problem the rest of the workflow fixes.
Stand up Portkey or LiteLLM (self-hosted if data residency matters), or use OpenRouter where you want a single bill across providers, and make it the only path to a model for application code — rotate the direct keys out of circulation after migration. The gateway gives you per-request metadata, retries and fallbacks, and the enforcement point for everything below. Platform credits in Copilot Studio and Agentforce cannot go through the gateway; they get their own budget line and alert in step 5.
Every request carries team, product or use case, environment, and an agent identifier. Reject untagged requests at the gateway in non-production immediately and in production after a two-week grace window. Export the tagged usage to Langfuse or LangSmith for trace-level cost analysis, so a cost spike can be followed to the exact prompt and tool loop that caused it rather than to a monthly total.
For each use case, name the default model, the fallback, and the conditions for escalating to a frontier model. Most classification, extraction, and routing work belongs on Haiku 4.5 at $1 input or Sonnet 5.5 and GPT-6.1 Sol at $2; reserve Opus 5.5 ($4/$20) and Fable 5.1 or GPT-6 Astra ($10/$50) for the tasks that measurably need them. Then structure prompts so stable system content and tool definitions are cache hits: at $0.20 to $0.25 per million for cached input, a long-context agent that re-reads the same 50,000 tokens of instructions every turn costs a fraction of the uncached version. Batch anything that is not latency-sensitive for the 50 percent discount both Anthropic and OpenAI offer.
Per team and per use case, set a monthly budget in the gateway with alerts at 50, 80, and 100 percent, and a hard stop for non-production. For Copilot Studio, compute daily credit burn from the published rates — Microsoft's own worked example is 7,200 credits a day for a 900-conversation support agent — and alert at 90 percent of prepaid packs, because enforcement at 125 percent disables agents and the alert must arrive while there is still time to buy a pack. That line applies to prepaid packs only — environments on the pay-as-you-go meter are billed for overage instead of disabled — and agent flows stop earlier: new flow runs are blocked once prepaid credits are fully consumed (100%), so an agent that depends on flows effectively has a 100% ceiling. For Managed Agents on the Claude platform, set per-session budgets (available since August 7, 2026) so a runaway loop ends itself.
Report cost per team and per use case, and the unit cost that the business understands: cost per resolved ticket, per document processed, per qualified lead. Showback for the first quarter, chargeback after. The unit cost is what lets a product owner decide whether an agent is worth running, and it is the number that stops the next shared-key experiment from becoming a surprise.
Known dates as of this writing: November 30, 2026 (Claude Sonnet 4.5 retires — migrate anything pinned to it), January 1, 2027 (Gemini Flash input and output prices double), and the end of the Gemini 4 Argon introductory period. Re-run the routing table two weeks before each, re-check cache hit rates monthly, and re-verify every vendor rate quarterly — the prices quoted here were read on October 3, 2026 and will not hold.
Use these templates as-is or customize for your business.
# Prices as read 2026-10-03 from vendor pricing pages. Re-verify quarterly.
# input/output are USD per 1M tokens; cache_read is the cached-input price.
models:
haiku-4.5: { provider: anthropic, input: 1.00, output: 5.00, cache_read: 0.10 }
sonnet-5.5: { provider: anthropic, input: 2.00, output: 10.00, cache_read: 0.20 }
opus-5.5: { provider: anthropic, input: 4.00, output: 20.00, cache_read: 0.20 }
fable-5.1: { provider: anthropic, input: 10.00, output: 50.00, cache_read: 0.25 }
gpt-6.1-sol: { provider: openai, input: 2.00, output: 10.00, cache_read: 0.10, long_context: { over_tokens: 272000, input: 4.00, cache_read: 0.20, output: 15.00 } }
gpt-6-astra: { provider: openai, input: 10.00, output: 50.00, cache_read: 1.00, long_context: { over_tokens: 272000, input: 20.00, cache_read: 2.00, output: 75.00 } }
gemini-4-argon: { provider: google, input: 2.00, output: 10.00, note: "intro; standard is 4.00 / 20.00" }
gemini-3.8-flash: { provider: google, input: 0.75, output: 3.75, note: "1.50 / 7.50 from 2027-01-01" }
routing:
classify_or_extract: { default: haiku-4.5, fallback: gemini-3.8-flash }
support_conversation: { default: sonnet-5.5, fallback: gpt-6.1-sol, escalate_to: opus-5.5, escalate_when: "two failed tool loops or explicit complexity flag" }
code_or_analysis: { default: opus-5.5, fallback: gpt-6.1-sol, escalate_to: fable-5.1, escalate_when: "human approval in ticket" }
batch_offline: { default: sonnet-5.5, mode: batch, discount: 0.5 }
required_tags: [team, use_case, env, agent_id]
untagged_requests: { nonprod: reject, prod: reject_after: 2026-11-01 }
budgets_usd_month:
- { team: support, use_case: tier1_agent, soft: 4000, alerts: [0.5, 0.8, 1.0], hard_stop: false }
- { team: finance, use_case: ap_extraction, soft: 900, alerts: [0.5, 0.8, 1.0], hard_stop: false }
- { team: "*", env: nonprod, soft: 500, alerts: [0.8, 1.0], hard_stop: true }
cache_policy:
min_stable_prefix_tokens: 2048 # put system prompt, tools, policy docs first
target_cache_hit_rate: 0.70 # alert below this
platform_credits: # cannot go through the gateway; tracked separately
copilot_studio: { pack_credits: 25000, pack_usd: 200, agents_disabled_at: 1.25, flows_blocked_at: 1.00, alert_at: 0.90, note: "prepaid packs only; pay-as-you-go bills overage instead" }
agentforce: { credits_per_100k_usd: 500, credits_per_action: 20, credits_per_voice_action: 30, alert_at: 0.90 }
calendar:
- { date: 2026-11-30, event: "Claude Sonnet 4.5 retires — migrate pins" }
- { date: 2027-01-01, event: "Gemini Flash prices double — re-run routing" }
- { date: TBD, event: "Gemini 4 Argon intro pricing ends" }Published consumption rates (learn.microsoft.com, page dated 2026-08-03): classic answer ........ 1 credit generative answer ..... 2 credits agent action .......... 5 credits tenant graph grounding 10 credits agent flow actions .... 13 credits per 100 voice ................. 10 / 35 / 75 credits per minute by tier Microsoft's own example: [(4 x 1) + (2 x 2)] x 900 customers = 7,200 credits/day YOUR AGENT conversations/day ............... ____ classic answers per conv ........ ____ x 1 generative answers per conv ..... ____ x 2 agent actions per conv .......... ____ x 5 graph groundings per conv ....... ____ x 10 voice minutes per conv .......... ____ x (10|35|75) = credits per conversation ...... ____ x conversations/day ............. ____ = daily burn x 30 ............................ ____ = monthly burn PACKS prepaid packs x 25,000 .......... ____ = prepaid credits ($200 per pack per month) vendor enforcement at 125% ...... prepaid x 1.25 = ____ (custom agents DISABLED here; prepaid packs only — pay-as-you-go bills overage instead) agent flow runs blocked at 100% of prepaid (not 125%) .. prepaid x 1.00 = ____ our alert at 90% ................ prepaid x 0.90 = ____ days of runway at current burn .. (prepaid - used) / daily burn = ____ RULE: if days of runway < time to procure a pack + 3 days, buy now.
PERIOD: ____________ BUSINESS UNIT: ____________ PREPARED BY: FinOps TOTAL AI SPEND $______ vs budget $______ (____%) Direct API via gateway $______ Cloud-hosted inference $______ Platform credits (Copilot Studio, Agentforce) $______ SaaS AI add-ons $______ BY USE CASE Spend Unit Units Cost/unit Prior month tier1_support_agent $_____ resolved _____ $_____ $_____ ap_extraction $_____ documents _____ $_____ $_____ sales_research_agent $_____ leads _____ $_____ $_____ EFFICIENCY Cache hit rate (target >= 70%) ........ ____% Share of tokens on frontier models .... ____% (routing policy target: ____%) Batch share of eligible workloads ..... ____% Untagged spend ........................ $_____ (target: $0) ALERTS FIRED THIS MONTH ................. ____ (list: team / use case / threshold) VENDOR ENFORCEMENT EVENTS ............... ____ (must be zero) UPCOMING PRICE EVENTS (next 90 days) .... [from calendar] ACTIONS .................................. 1. ____ 2. ____ 3. ____
Get a new AI workflow every week. Prompts, tool stacks, and ROI math included.
Where this workflow tends to break in production — and what to put in place before you ship it.
Shared API keys hide which team caused a spend spike
Mitigation: All direct traffic through the gateway with mandatory tags; untagged requests rejected; direct keys rotated out after migration.
Vendor enforcement disables agents before anyone is alerted
Mitigation: Alerts at 90% of prepaid platform credits, below both the 125% agent cutoff and the 100% agent-flow block; days-of-runway computed weekly.
Cheaper model routed in without quality measurement
Mitigation: Routing changes require an evaluation run against the golden dataset; escalation rules written per use case.
Prompt structure defeats caching and the cache lever never materializes
Mitigation: Stable prefix placed first; cache hit rate tracked with a 70% floor and an alert below it.
Price cliff or model retirement arrives unnoticed
Mitigation: Pricing calendar in the policy file; routing table re-run two weeks before each dated event; rates re-verified quarterly.
Skip the gateway if you run one or two agents on one provider with one owner — a budget alert in the vendor console and a monthly look at the invoice is proportionate, and a gateway adds a dependency you then have to run. Do not route for cost before you have an evaluation harness: moving a workload to a cheaper model without measuring quality is how you save $300 and lose a customer, and the agent evaluation workflow in this library is the prerequisite. Do not treat the prices in this page as current — they were read on October 3, 2026 and vendors changed them several times this year alone. And do not let FinOps become the team that says no to experiments; the point of tagging and showback is that teams can see what they are spending and decide for themselves, with a hard stop only in non-production.
A phased approach to get this workflow running and delivering ROI.
Days 1–30
Foundation
Days 31–60
Optimization
Days 61–90
Scale
Salesforce bills actions or conversations, Microsoft bills credits and switches you off at 125%, Google bills seats, OpenAI bills tokens plus a sandbox, Anthropic bills tokens plus session-hours. One 900-conversation-a-day agent, priced on each where the published numbers allow.
Half the dental AI vendors marketing "HIPAA compliance" cite a Security Rule update that has not been finalized. Here is what to ask for in writing, and the one design choice that makes most of the problem disappear.
Per-minute platform pricing makes building look cheap. It is cheap. The platform was never the expensive part — and once you see the real cost line, the build-or-buy question answers itself.
One practical AI workflow per week. No fluff.
Get the full guide with step-by-step setup, workflow templates, and copy-paste assets.