The September 2026 price war dropped the floor of the LLM API market to one tenth of a cent per output token, the frontier tier to a fifth of its January price, and the mid-tier to roughly the same dollars-per-million as a small coffee. Every figure below was pulled from each vendor’s own pricing page on 2026-09-28; the only secondary is Simon Willison’s Opus and Sol and Luna post, and it is cited only for the price-change narrative, not the numbers themselves.
TL;DR
- New floor: OpenAI GPT-6 Luna at $0.10/$0.50 per million tokens is one tenth the price of Anthropic Haiku 4.5 ($1/$5) and one fifth the price of GPT-4.1 Nano ($0.10/$0.40, April 2025).
- New mid-tier: GPT-6 Sol at $2/$10, Grok 4.7 at $2/$6, and Claude Sonnet 5 at $2/$10 now sit at the same dollar-per-million; the difference is output quality, not sticker price.
- Frontier: Claude Opus 5.5 at $4/$20 is a 20 percent cut from Opus 4.x ($5/$25). Claude Fable 5.1 at $10/$50 and GPT-6 Astra at $10/$50 sit on top.
- Cache read cuts: Anthropic dropped the cache-read multiplier from 10 percent to 5 percent on Opus 5.5 (and to 2.5 percent on Fable 5.1). Long system prompts are now between five and forty times cheaper to re-read.
- November 2026 cliff: OpenAI’s promotional pricing on GPT-5.6 ends on 2026-11-21, after which a 25 percent price increase is scheduled. Lock workloads that depend on GPT-5.6 into GPT-6 Sol before the cliff.
- Local still wins at volume. Per the cost-crossover article, self-hosted Qwen3 and Llama 3.3 cross below any API price once monthly token volume passes a threshold that depends on hardware, but not on sticker price.
What “after the September price war” actually means
Three independent price moves landed in the same two-week window in September 2026. Anthropic cut Opus 5.5 to $4 per million input tokens and $20 per million output tokens, a 20 percent drop from the $5/$25 that Opus 4.5, 4.6, 4.7, 4.8, and the original Opus 5 had shared since launch. OpenAI released GPT-6 with three tiers (Sol at $2/$10, Luna at $0.10/$0.50, and Astra at $10/$50), and Simon Willison’s price table confirms the September positioning as the new floor for the API market.
The most under-reported move is the Anthropic cache-read reduction. A cache hit on Opus 5.5 costs $0.20 per million tokens, which is 5 percent of the base input price rather than the 10 percent that every prior Claude model used. On Fable 5.1 the multiplier drops to 2.5 percent, so a cache hit costs $0.25 per million tokens on a $10 base. For long system prompts and retrieval-augmented chat, that is the difference between cache saving money and cache paying for itself on a single read.
The new floor: GPT-6 Luna versus Haiku 4.5
The cheapest current API model on each vendor’s pricing page is a useful anchor. On 2026-09-28 the anchors are:
| Model | Input $/MTok | Output $/MTok | Cached input $/MTok | Source |
|---|---|---|---|---|
| Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 | Anthropic pricing |
| OpenAI GPT-6 Luna | $0.10 | $0.50 | $0.01 | OpenAI pricing |
Luna is ten times cheaper than Haiku 4.5 on input and on output, and ten times cheaper than Haiku’s own cache read. Simon Willison notes that GPT-4.1 Nano in April 2025 was $0.10/$0.40 and GPT-5 Nano in August 2025 was $0.05/$0.40; GPT-6 Luna at $0.10/$0.50 matches GPT-4.1 Nano on input and raises the output price 25 percent above both prior Nanos. For bulk classification, log triage, embedding-style scoring, and “summarise this transcript” workloads, the new floor is cheap enough that the choice is no longer whether to send the request but how many workers to fan it out across.
The mid-tier sweet spot: GPT-6 Sol, Grok 4.7, Sonnet 5
Three models from three vendors now share the $2 input price point, which makes the mid-tier a quality comparison rather than a price comparison. From the vendor pricing pages on 2026-09-28:
| Model | Input $/MTok | Output $/MTok | Cached input $/MTok | Source |
|---|---|---|---|---|
| Claude Sonnet 5 | $2.00 | $10.00 | $0.20 | Anthropic pricing |
| OpenAI GPT-6 Sol | $2.00 | $10.00 | $0.20 | OpenAI pricing |
| xAI Grok 4.7 (under 200k) | $2.00 | $6.00 | $0.50 | xAI pricing |
Grok 4.7 is the cheapest on output at $6 per million tokens; the others are at $10. The reverse is true on cache reads: Sonnet 5 and GPT-6 Sol are at 10 percent of input ($0.20), but Grok 4.7 is at 25 percent ($0.50). For long-context retrieval, Sonnet 5 and GPT-6 Sol are the cheaper choice on cache, while Grok 4.7 wins on raw output for short, non-cached calls.
Grok 4.7 also has a two-tier context price: $4 input, $12 output, and $1.00 cached at the 200k token threshold. The other vendors do not currently publish a context-tier break.
The frontier tier: Opus 5.5 versus Fable 5.1 versus GPT-6 Astra
Three models currently sit at the top of each vendor’s catalogue. From the vendor pricing pages on 2026-09-28:
| Model | Input $/MTok | Output $/MTok | Cached input $/MTok | Source |
|---|---|---|---|---|
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 | Anthropic pricing |
| Claude Fable 5.1 | $10.00 | $50.00 | $0.25 | Anthropic pricing |
| OpenAI GPT-6 Astra | $10.00 | $50.00 | $1.00 | OpenAI pricing |
Opus 5.5 at $4/$20 is the cheap frontier. Fable 5.1 and GPT-6 Astra are the same price on input and output ($10/$50) but Fable 5.1 has a 2.5 percent cache read multiplier ($0.25), while GPT-6 Astra is at 10 percent ($1.00). Anthropic also publishes a fast-mode premium for Opus 5.5 of $8/$40 per million, doubling the input and output rates for lower-latency inference.
The capability story is not in this article; the capability ceiling piece covers benchmarks, licence terms, and context windows for the same model set, refreshed on 2026-09-27.
Cached input is the real discount
The price war is more interesting in cache than in sticker. From the Anthropic pricing page:
- Opus 5.5: cache hit at 5 percent of base ($0.20 per million, against a $4 base).
- Fable 5.1: cache hit at 2.5 percent of base ($0.25 per million, against a $10 base).
- All other Anthropic models: cache hit at 10 percent of base ($1 per million on Fable 5, $0.50 per million on Opus 5, and so on).
Simon Willison’s post calls out the 60 percent drop in cache-read pricing as the single biggest September move. For a 50,000-token system prompt with a 1-hour cache on Opus 5.5, the prompt costs about $0.40 to write (50,000 tokens at $8/MTok for the 2x 1-hour write multiplier) and $0.01 to read on every subsequent request (50,000 tokens at $0.20/MTok). At a thousand re-reads the prompt amortises to about one cent per call.
The OpenAI side is less dramatic: GPT-6 Sol and Luna both price cache reads at 10 percent of base, the same multiplier as prior GPT generations. The relative gain there is the underlying base-price drop, not a cache-multiplier change.
Where local inference still beats the API
The cost crossover has shifted since August. Per the cost-crossover article, the monthly token volume at which a self-hosted model becomes cheaper than the cheapest API option depends on three numbers: the amortised hardware cost, the electricity draw, and the inference throughput. After the September cuts the API floor dropped from $1/$5 (Haiku 4.5) to $0.10/$0.50 (GPT-6 Luna), so the crossover point moved up.
Two workloads still favour local inference regardless of sticker price:
-
Workloads that exceed a fixed monthly volume. A self-hosted
qwen3:32bon a 24 GB card at roughly 30 tokens per second runs roughly 78 million tokens per month continuously. At the GPT-6 Luna price of $0.10/$0.50, that volume costs $7.80 input and $39 output per month on the API, far below the electricity bill of a 24/7 GPU. Volume moves the needle in local’s favour. -
Workloads that cannot leave the machine. Per the privacy hub, prompts that contain code, customer data, or unreleased product information should not flow through a hosted endpoint by default. Local inference with a Qwen3 or Llama 3.3 open-weight model keeps the bytes on the box.
For workloads below the crossover that are safe to send to the API, GPT-6 Luna is the new default. For workloads above the crossover or that must stay local, the best-local-LLM guide and the cost-crossover piece cover the hardware ladder in detail.
The November 2026 cliff
Simon Willison flagged on 2026-09-22 that GPT-5.6 has a scheduled 25 percent price increase for November. The OpenAI pricing page confirms that GPT-5.6 promotional pricing runs through 2026-11-21 without specifying the post-November rate. Simon Willison’s GPT-5.6 Sol row lists $4/$20 today, which is exactly the new Opus 5.5 rate; a 25 percent rise would put GPT-5.6 Sol at $5/$25, the same number Opus 4.x carried for two years.
Two operational notes:
- Migration window: any production workload on GPT-5.6 has roughly eight weeks to either accept the 25 percent rise or migrate to GPT-6 Sol ($2/$10 today, no scheduled change).
- Lock-in risk: the November date is the promotional end, not a guaranteed deprecation. Plan as if GPT-5.6 Sol will exist at $5/$25 from 2026-11-22, but verify on the vendor pricing page before committing.
The Bottom Line
After the September 2026 cuts the API market has three usable tiers and one sub-tier: cheap (GPT-6 Luna at $0.10/$0.50), mid (Sonnet 5, GPT-6 Sol, Grok 4.7 at $2 with output from $6 to $10), and frontier (Opus 5.5 at $4/$20, Fable 5.1 and GPT-6 Astra at $10/$50). The cache-read multiplier is where the real saving lives on long-system-prompt workloads, and the November 21, 2026 GPT-5.6 cliff is the date to plan around. Local inference still wins on volume or on workloads that must stay on the machine; for everything else, GPT-6 Luna is the new default pick and GPT-6 Sol is the mid-tier pick unless output quality benchmarks point elsewhere.
Related on Intelligibberish
- Leading Frontier AI Models: Capability Ceiling (September 2026) - the capability ceiling, refreshed 2026-09-27, with the same models at the same vendor pricing.
- When Local AI Beats the API: The Cost Crossover (August 2026) - the monthly volume at which a self-hosted model crosses below the API, with the hardware and electricity assumptions.
- Best Local LLM for Everyday Tasks - the open-weight ladder from 8 GB laptops to 32 GB workstations.