WorkflowStack AI
WorkflowsIndustriesToolsGuidesAI QuizBlogEnterprise
Get Free Workflows
WorkflowStack AI

Practical AI workflows for SMB operators and enterprise teams. No fluff. No hype. Just what ships.

Library

  • All Workflows
  • Industries
  • Enterprise
  • Tools
  • Guides

Company

  • About
  • Blog
  • Newsletter
  • Contact

Stay Updated

Weekly workflow ideas for operators and enterprise teams.

Get Free Workflows →

© 2026 Blueteem LLC. All rights reserved.

Privacy PolicyTerms of Service
HomeBlogThe October 2026 Model Price Sheet: Astra, Fable 5.1, Opus 5.5, GPT-6.1 Sol, Sonnet 5.5 and Gemini 4 Argon, With a 10,000-Ticket Bill
September 30, 2026

The October 2026 Model Price Sheet: Astra, Fable 5.1, Opus 5.5, GPT-6.1 Sol, Sonnet 5.5 and Gemini 4 Argon, With a 10,000-Ticket Bill

Two labs now sell the workhorse tier at $2 in and $10 out, and Google has announced the same price for Gemini 4 Argon, which is not yet available to developers. The differences that remain are cache-read price, long-context surcharges, and two dated cliffs: Gemini Flash doubles on 1 January 2027 and Sonnet 4.5 retires on 30 November 2026.

The one-paragraph version

If you are building a support or back-office agent and have not chosen a model, the mid tier is a near-tie on list price: Claude Sonnet 5.5 and GPT-6.1 Sol are $2 per million input tokens and $10 per million output; Google has announced Gemini 4 Argon at the same introductory price, but it is not yet sold to API customers. Pick on quality for your task and on three second-order prices: what a cached input token costs, whether long prompts carry a surcharge, and whether the price you see is introductory. Then put two dates in the calendar: 30 November 2026 (Claude Sonnet 4.5 retires) and 1 January 2027 (Gemini Flash prices double).

The sheet

Prices per million tokens. Prices and plan names are as published on each vendor's own pricing page; last verified 3 October 2026. "Cache read" is the price of an input token served from the provider's prompt cache.

The sheet
ModelInputCache readOutputNotes
GPT-6 Astra (OpenAI)$10.00$1.00$50.00Above 272K input tokens: $20.00 / $2.00 / $75.00. Batch and Flex 50% off. Fast $20 / $2 / $100. Ultrafast $60 / $6 / $300, Astra only
Claude Fable 5.1 (Anthropic)$10$0.25$50Cache read is 0.025x input. 1M context at standard pricing. Batch $5 / $25
Claude Opus 5.5$4$0.20$20Cache read 0.05x. Fast mode $8 / $40. Batch $2 / $10. Opus 5 remains $5 / $25
Gemini 4 Argon (Google)$2 introductory95% off input$10 introductoryAnnounced 30 September 2026 at introductory pricing; available at launch only to Fairwind Program testers, not to API customers; not on the Gemini API pricing page. Standard price after the intro period: $4 / $20. 1M-token output limit
GPT-6.1 Sol$2.00$0.10$10.00Batch and Flex 50% off
GPT-6 Sol$2.00$0.20$10.00
Claude Sonnet 5.5$2$0.20$10Launched 28 September 2026. Batch $1 / $5
Claude Sonnet 5$2$0.20$10The planned rise to $3 / $15 on 1 September "will not occur" (Anthropic, 10 August 2026)
Claude Sonnet 4.5$3$0.30$15Deprecated 30 September 2026; retirement on the Claude API "scheduled for November 30, 2026"
Claude Haiku 4.5$1$0.10$5Batch $0.50 / $2.50
Gemini 3.8 Flash / 3.7 Flash$0.75 through 31 December 2026; $1.50 from 1 January 2027$0.075 through 31 December 2026; $0.15 from 1 January 2027$3.75 through 31 December 2026; $7.50 from 1 January 2027The step-up is printed on the pricing page itself
GPT-6 Luna$0.10$0.01$0.50

Three structural notes that matter more than any single row:

Cache reads are where the labs actually compete. A support agent sends the same system prompt and knowledge snippets on every turn. Anthropic prices a Fable 5.1 cache hit at 2.5% of input and an Opus 5.5 hit at 5%; OpenAI's GPT-6.1 Sol hit is 5% and Astra's is 10%; Google has announced 95% off for Argon. On a heavily cached workload these ratios move the bill more than the headline input price does.

Long context is a surcharge at OpenAI and not at Anthropic. OpenAI's page defines long context as more than 272K input tokens and roughly doubles the rate. Anthropic's page says Claude 4.6 and later include "the full 1M token context window at standard pricing." If your agent reads whole documents, that is a real difference.

Tokenizers differ. Anthropic states that Claude 4.7 and later use a tokenizer that "produces approximately 30% more tokens for the same text" than Claude Sonnet 4.6 and earlier. If you are migrating off Sonnet 4.5 and comparing against your current token counts, multiply the Anthropic rows by roughly 1.3 before you compare. Per-token prices across labs are not directly comparable without measuring your own text.

Get one new AI workflow every week

Practical playbooks, prompts, and tool stacks — for the SMB operators and enterprise teams actually shipping AI.

The worked bill: 10,000 support tickets a month

Assumptions, stated so you can change them: each ticket is one model call with 3,000 input tokens, of which 2,000 are a cached system prompt and knowledge block and 1,000 are fresh (the customer message and retrieved context), plus 700 output tokens. That is 10 million uncached input, 20 million cached input and 7 million output tokens a month. Cache-write charges are ignored (they recur once per cache window, not per ticket). No Batch discount, because support is interactive. Anthropic's own worked example on the same page lands at "~$37.00 per 10,000 tickets" on Haiku 4.5 with a different token mix; the point of the table is the ratios, not the cents.

The worked bill: 10,000 support tickets a month
ModelMonthly bill for 10,000 tickets
GPT-6 Luna$4.70
Gemini 3.8 Flash, cached, until 31 Dec 2026$35.25
Claude Haiku 4.5$47
Gemini 3.8 Flash, no caching, until 31 Dec 2026$48.75
Gemini 3.8 Flash, cached, from 1 Jan 2027$70.50
GPT-6.1 Sol$92
Gemini 4 Argon, introductory (announced price; model not yet purchasable)$92
Claude Sonnet 5.5 / Sonnet 5 / GPT-6 Sol$94
Gemini 3.8 Flash, no caching, from 1 Jan 2027$97.50
Claude Opus 5.5$184
Gemini 4 Argon, standard price (announced price; model not yet purchasable)$184
Claude Fable 5.1$455
GPT-6 Astra$470

Arithmetic for one row, so you can check the rest: Sonnet 5.5 is 10M × $2 + 20M × $0.20 + 7M × $10, divided by one million, which is $20 + $4 + $70 = $94.

What the table says:

  • At the mid tier the two models you can actually buy are within $2 of each other on this workload, and Argon is announced at the same point. Model choice there is a quality decision, not a price decision.
  • The Flash cliff is a 2x on the exact same traffic. A team on 3.8 Flash at $35.25 a month with caching ($48.75 without) wakes up on 1 January at $70.50 ($97.50 without), which is close to Sonnet 5.5 money without the model change.
  • When Argon reaches API customers, its introductory price buys Sonnet-class economics and its standard price is Opus 5.5 economics. Google's post gives neither a developer availability date nor the end of the introductory period. Budget at $184, enjoy $92.
  • The frontier tier is five times the mid tier. Astra and Fable 5.1 are for the tickets your mid-tier model gets wrong, routed to them deliberately, not for every ticket.

The two dates

30 November 2026. Anthropic announced on 30 September that Claude Sonnet 4.5 is deprecated, with retirement on the Claude API scheduled for 30 November 2026. Anything pinned to that model string returns errors after that date (Anthropic's August retirement of Opus 4.1 worked the same way). The migration is to Sonnet 5.5 at a lower list price, with the 30% tokenizer caveat above, and it needs an eval run, not a string replace. The agent evaluation and regression harness is the workflow for exactly that.

1 January 2027. Gemini 3.7 and 3.8 Flash go from $0.75 to $1.50 input and $3.75 to $7.50 output. That is printed on Google's pricing page. If Flash is your default cheap tier, decide before December whether the January price still beats Haiku 4.5 at $1 / $5 or GPT-6 Luna at $0.10 / $0.50 for your task.

Two things the sheet does not show

Platform fees. The sheet is the token price only. If you run the model inside a hosted agent runtime, add the runtime's own meter. Anthropic's Claude Managed Agents adds $0.08 per session-hour while a session is running; OpenAI's hosted containers run $0.03 to $1.92 per 20-minute session depending on container size. Our enterprise agent platform comparison works those through.

Regional inference. Anthropic charges 1.1x for US-only inference on Claude 4.6 and later. Bedrock and Google Cloud regional endpoints carry a 10% premium over global endpoints. If data residency is a requirement, it is also a line item.

Next step

Price your own workload, not mine: pull a week of real tickets, count tokens with each provider's counting endpoint, and rerun the table. Then build the routing so the expensive model only sees the tickets that need it; the RAG support agent workflow and support inbox triage are the two halves of that. For the surrounding decisions, read how AI agent pricing models work and the free AI Agents Starter Kit; enterprise teams should go to the Enterprise AI Agent Platform Buyer's Guide. Head-to-heads: Claude API vs. OpenAI API, Claude API vs. Gemini, Gemini vs. OpenAI API. Tool pages: OpenAI API, Claude API, Gemini. For the enterprise view of the same decision, start at the enterprise hub.

Related Workflows

Multi-Agent Customer Support Deflection at Scale

A tiered agent system that resolves routine support tickets end-to-end, escalates the rest with full context, and keeps humans on the hard 20%.

View workflow

RAG Customer Support Agent

A retrieval-grounded support agent that answers tier-1 tickets from your docs and ticket history — escalates the rest with full context.

View workflow

Knowledge Base & Help Center Builder

Build a self-service knowledge base powered by AI search that deflects 40-60% of support tickets.

View workflow

Keep Reading

September 27, 2026

Fin vs. Zendesk vs. Gorgias vs. Freshdesk: The AI Support Agent Is Now Priced Four Different Ways

Fin bills an outcome, Zendesk a resolution, Gorgias an interaction, Freshdesk a 72-hour session, HubSpot a resolved conversation. Same 3,000 tickets, five invoices that differ by more than 3x.

September 1, 2026

The AI Receptionist Decision for Law Firms: Smith.ai vs. Ruby vs. Rosie vs. Building It Yourself

Smith.ai will sell you the same intake call two ways — $2.00 handled by AI or about $10 handled by a human. Knowing which calls belong in which bucket is the whole decision.

October 1, 2026

Agentforce vs. Copilot Studio vs. Gemini Enterprise vs. OpenAI Agents API vs. Claude Managed Agents: Five Billing Units for One Enterprise Agent

Salesforce bills actions or conversations, Microsoft bills credits and switches you off at 125%, Google bills seats, OpenAI bills tokens plus a sandbox, Anthropic bills tokens plus session-hours. One 900-conversation-a-day agent, priced on each where the published numbers allow.

Found this helpful?

Get weekly AI workflow ideas in your inbox.