Leading Frontier AI Models: The Capability Ceiling (August 2026)

What sits at the top of the field right now, hosted or downloadable, with every licence, price and index score re-read from a primary source today.

Updated August 9, 2026

This page answers one question: what sits at the capability ceiling right now, regardless of licence, price, or whether the weights can be downloaded at all. It is the opening entry in a monthly series, and the ranking will be rebuilt from primary sources each month rather than edited in place.

Two neighbouring pages answer different questions, and for most readers of this site they are the better starting point. The open-weight ranking covers only models whose weights can be downloaded, and treats the licence as a first-class constraint on what you may then do with them. The VRAM tier index maps card memory to a model and quantization that will actually load. This page sets both constraints aside on purpose. Most of what follows will not run on consumer hardware, and one entry cannot be bought at any price without an invitation.

How this ranking was built

Every hosted model below was read from its vendor’s own model or pricing page on 9 August 2026. Every open-weight entry was read from the Hugging Face model API the same day: cardData.license and license_name for the licence, gated for the access flag, and the byte-exact safetensors.total for the parameter count. No figure here is quoted from recollection.

The ranking spine is the Artificial Analysis Intelligence Index, a third-party composite rather than a vendor’s own scorecard. The version is load-bearing. The methodology page states v4.1.1, built from nine evaluations weighted across agents (34 percent), coding (24 percent), scientific reasoning (24 percent) and general tasks (18 percent), including GDPval-AA v2, Terminal-Bench v2.1, SciCode, GPQA Diamond and Humanity’s Last Exam. The leaderboard table itself does not print a version string, so that version number comes from the methodology page and not from the table. Scores produced under different index versions sit on different scales and cannot be compared with each other.

The leaderboard lists a separate row for each reasoning-effort setting, so one model appears several times at different scores. The table below takes each model’s highest score among the rows read today.

The field, August 2026

Index (v4.1.1)ModelMade byWeightsLicence or access modelBest at
63Claude Opus 5AnthropicHostedClaude API, Amazon Bedrock, Google Cloud, Microsoft Foundry; $5 / $25 per MTokComplex agentic coding and enterprise work
62Claude Fable 5AnthropicHostedSame surfaces, generally available since 9 June 2026; $10 / $50 per MTokLong-running agents at the highest generally available capability
61GPT-5.6 SolOpenAIHostedOpenAI API; $5 / $30 per MTokComplex professional work across a 1.05M-token context
60Kimi K3Moonshot AIOpenCustom kimi-k3 licence, ungatedThe only downloadable model inside the index top ten
58Qwen3.8 MaxAlibabaHostedNot listed on Alibaba’s own Model Studio page read todayAccess route not verifiable from a vendor source today
57Muse Spark 1.2MetaNOT VERIFIEDNOT VERIFIEDNOT VERIFIED
57GPT-5.6 TerraOpenAIHostedOpenAI API; $2 / $12 per MTokBalancing intelligence against cost
56Grok 4.5xAIHostedxAI API; $2 / $6 per MTok below 200k tokens, $4 / $12 at or aboveA 500k context at flagship quality
55Claude Sonnet 5AnthropicHosted$3 / $15 per MTok, introductory $2 / $10 through 31 August 2026Speed with near-top intelligence
Not scoredClaude Mythos 5AnthropicHostedInvitation only, Project Glasswing; no self-serve sign-upDefensive cybersecurity work, per Anthropic’s own framing
Not scoredDeepSeek-V4-ProDeepSeekOpenMIT, ungated1.60T parameters under a plain permissive licence
Not scoredInklingThinking MachinesOpenApache-2.0, ungatedThe largest Apache-2.0 model found today; text, image and audio input
Not scoredGLM-5.2Z.aiOpenMIT, ungatedPermissive frontier weights at 753B
Not scoredK-EXAONE 2.0 750B A37BLG AI ResearchOpenApache-2.0, ungatedMultilingual work and long-context retrieval
Not scoredGemini 3.5 FlashGoogleHostedGemini API; $1.50 / $9.00 per MTokSustained agentic and coding performance, per Google’s own wording

The rows marked “Not scored” are grouped rather than ranked. No index score for any of them appeared in the top fifteen read today, and inventing a position for them would be exactly the error this page exists to avoid. Their placement reflects nothing except that they were not scored in the table consulted.

The ceiling is rented, not owned

Fourteen of the fifteen index positions read today belong to models whose weights could not be found on Hugging Face today. The single exception is Kimi K3, which reports 2,779,931,837,184 parameters in safetensors.total and is ungated on Hugging Face under a custom licence named kimi-k3 in its card data. Anyone may download it. Almost nobody can serve it: at 2.78T total parameters it is a multi-node deployment, not a workstation one.

The top of the market is a rental market, and its one downloadable entry is downloadable in the way a container ship is purchasable.

The licence still buys something real even when the hardware does not. Weights under MIT or Apache-2.0 can be inspected, fine-tuned by anyone with a cluster, and served by a provider other than the lab that trained them. DeepSeek-V4-Pro returns mit with gated: false; Inkling returns apache-2.0 with a licence link pointing at apache.org. Those are portability guarantees, not self-hosting guarantees, and the two are routinely confused.

One entry sits outside both categories. Anthropic’s documentation states that Claude Mythos 5 shares Claude Fable 5’s specifications and pricing but is offered in limited availability to approved customers in Project Glasswing, invitation only, with no self-serve sign-up. A model whose capability cannot be independently measured because it cannot be independently purchased is a real feature of the current field, and no public leaderboard covers it.

What this means if you came here for local AI

Nothing on the scored half of this table will run on a graphics card you own. That is not a reason to ignore it, and it is not a reason to buy hardware.

The gap between the ceiling and what fits in 24GB is one you close with task selection, not with money. A model scoring in the fifties on this composite is being measured on agentic coding, scientific reasoning and long-context retrieval. Summarisation, extraction, classification, rewriting and retrieval-augmented answering are not what separates these systems, and those are the jobs most self-hosted deployments actually run.

The recommendation is therefore a split one. Use the VRAM tier index to pick something that loads on your card and handles routine work locally, where the data never leaves the machine. Keep an API key for the few tasks where the ceiling genuinely matters, and read the open-weight ranking before putting any downloadable model into production, because its licence column is where the expensive surprises live.

Price is a useful check on how much of the ceiling you need. Claude Sonnet 5 sits at $3 / $15 per MTok against Claude Fable 5 at $10 / $50, and the index separates them by seven points. GPT-5.6 Luna is listed at $0.20 / $1.20. The curve from adequate to best is steep at the top and flat for a long way below it.

What was searched for and not found

Absences here are stated with the exact query behind them.

Qwen3.8 Max weights. The Artificial Analysis table lists Qwen3.8 Max at 58. The Hugging Face API returns HTTP 401 with {"error":"Invalid username or password."} for Qwen/Qwen3.8-Max, Qwen/Qwen3.8, Qwen/Qwen3.8-27B, Qwen/Qwen3.8-397B-A17B and Qwen/Qwen3.8-Max-Instruct. That is the same response a control path known not to exist returns, so it indicates no publicly visible repository rather than a gate. The Qwen account sorted by creation date shows nothing newer than Qwen/Qwen3-ForcedAligner-0.6B-hf from 26 June 2026. Alibaba’s own Model Studio list, read today, names qwen3.7-max as its most capable text model and contains no 3.8 entry at all. Third-party repositories named Qwen3.8-27B were created on 5 and 6 August by accounts unaffiliated with Alibaba; their provenance is unverified and they are not evidence of an official release.

Muse Spark 1.2. This model appears in the leaderboard at 57, attributed to Meta, and could not be corroborated against any Meta source. A Hugging Face search for Muse-Spark returns nothing; the meta-llama account has published nothing newer than 28 April 2025; the facebook account’s recent uploads are vision and reconstruction models. Meta’s pages at ai.meta.com and llama.com did not return a readable model listing to an automated request. The row is marked NOT VERIFIED rather than described.

Frontier weights from Anthropic and OpenAI. No Anthropic model account on Hugging Face returns repositories. The openai account’s most recent language models remain gpt-oss-120b and gpt-oss-20b from 4 August 2025. meta-llama/Llama-5 returns the same not-found response as the control. The xai-org account tops out at grok-2 from 22 August 2025, so Grok 4.5 is hosted only.

A Gemini entry in the top fifteen. None appeared in the rows read today. Google’s documentation describes gemini-3.5-flash as its “Most intelligent model for sustained frontier performance on agentic and coding tasks”, while listing gemini-3.1-pro-preview as a preview. Absence from one third-party table on one day is weak evidence and is recorded as such.

A parameter-count discrepancy. Inkling’s model card states “975B total, 41B active”, while the API reports 952,377,623,626 in safetensors.total. Both figures are given rather than picking one.

All figures read from vendor documentation and the Hugging Face model API on 9 August 2026. Benchmark positions come from a third-party index whose version is recorded above; vendor capability claims quoted here are self-reported. Prices and model lineups change weekly.