OpenRouter Rankings July 2026: Who's Actually Winning the AI Model Race

About 20 min read · MACCOME · Last updated: July 27, 2026

Who this is for: developers and engineering leads choosing models through OpenRouter rankings, not benchmark slides. Bottom line (data through July 25, 2026): Xiaomi Mimo V2.5 tops daily token volume at 1.4T/day; Chinese labs hold about 46% of platform share (up from under 2% a year ago); US labs fell from ~70% to 30-36%. What you get: Top 12 model table, vendor and app leaderboards, pricing comparison, usage-vs-quality barbell analysis, August outlook, a six-step tiered-routing runbook, Python routing example, and FAQ. Builds on our June OpenRouter rankings and pairs with the OpenRouter API guide.

insights

TL;DR — 30-second verdict

  • Daily #1 (7/25): Mimo V2.5 (Xiaomi) at 1.4T tokens/day. DeepSeek V4 Flash second (943.9B). Seven of the top twelve models are Chinese-origin.
  • Stable provider #1: DeepSeek still leads weekly vendor share at 16-18%, but the "model of the month" crown keeps rotating inside the Chinese camp.
  • Barbell market: Cheap open models own volume; Claude and GPT still own hard-task spend. Opus 5 (7/24) defends the premium side at $5/$25 per 1M tokens.
  • Hidden half: Hermes Agent (~45% app share) plus roleplay apps move enormous volume that enterprise AI coverage rarely mentions.

Six Pain Points: What Teams Get Wrong Reading OpenRouter Rankings

  1. Treating daily #1 as a durable quality crown. Mimo V2.5 leads on July 25, but the monthly model winner has rotated every quarter (MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 now). Rankings shift daily; provider stability matters more than a snapshot.
  2. Equating token volume with "best LLM". OpenRouter rankings explained in one line: they measure paid production traffic, not benchmark scores. A cheap model behind a high-traffic consumer app can outrank a smarter model teams reserve for the hardest 10% of work.
  3. Ignoring the 46% vs 61% methodology gap. Chinese vendors labeled in the top ranks sum to about 46%. Broader English-language cuts that include more China-origin traffic report 61%. Direction is the same; do not overfit one denominator.
  4. Missing the spend barbell on hard tasks. General chat (35.7%), agentic workflows (30.4%), and code (26.5%) dominate spend by category. On classification and complex reasoning, Claude Sonnet 4.6 and Claude Opus 4.7 tie at 13.5% of spend each; GPT-5.5 holds 11.6%. Volume leaders barely register there.
  5. Looking at models but not apps. Hermes Agent alone is ~45% of tracked app tokens. Coding agents (Kilo Code, OpenClaw, Claude Code) and roleplay apps (Janitor AI, ISEKAI ZERO) explain why certain cheap models spike — context our Opus 5 / Kimi K3 week-in-review also touched from the compliance angle.
  6. Routing everything to one leaderboard winner. DeepSeek V4 Flash is the cheapest LLM API for coding at scale, but critical agent steps still need Opus 5 or GLM 5.2. Teams without tiered routing pay either too much or fail too often.

July 2026 Model Rankings: Top 12 by Daily Token Volume

Figures below are from OpenRouter as of July 25, 2026. Rankings update daily — verify before citing in production docs.

RankModelVendorDaily tokens30-day total
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra 550B (free)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot AI157.6B1.6T (fastest rise)
10Ling 3.0 FlashInclusionAI (Ant)128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

Seven of twelve are Chinese-origin. US models still present via NVIDIA's free Nemotron tier, Gemini 3 Flash, and Claude Sonnet 5. Note the volatility: Claude Opus 4.8 ranked inside the top 10 on July 24 data but dropped out of the top 12 by July 25 as Ling 3.0 Flash climbed — proof that weekly reviews beat one-off screenshots.

Vendor share (~7-day window, blended sources)

VendorOriginToken share (approx.)
DeepSeekChina16-18% (most stable #1)
XiaomiChina8-18% (Mimo V2.5 spike)
AnthropicU.S.10-15%
TencentChina8-13%
GoogleU.S.8-13%
Z.aiChina4-7%
OpenAIU.S.6-8%
NVIDIAU.S.~5%
MiniMaxChina4-8%
Moonshot AIChina3-4%
Alibaba QwenChina1-4%

Chinese AI models market share combined: ~46%. US labs (OpenAI + Anthropic + Google) sit at roughly 30-36%, down from about 70% in mid-2025. This is pricing math: DeepSeek V4 Flash input runs ~$0.05-$0.14 per million tokens versus GPT-5.5 at ~$5 — a ~35x gap. When open-weight models are good enough for bulk work, developers route real money accordingly.

Usage vs Quality: The Barbell Nobody Puts in Headlines

OpenRouter rankings July 2026 tell you who moves tokens. They do not tell you who wins hard tasks. Spend-by-category data paints a dumbbell:

  • High-volume, error-tolerant workloads (chat, creative writing, roleplay, routine coding): Chinese open models — Mimo V2.5, DeepSeek V4 Flash, Hy3, GLM 5.2.
  • Low-volume, low-error-tolerance workloads (classification, complex agent planning, regulated reasoning): Claude Sonnet 4.6 and Opus 4.7 at 13.5% of classification spend each; GPT-5.5 at 11.6%.

Anthropic's July 24 Claude Opus 5 launch reinforces the premium side: FrontierBench v0.1 at 43.3% versus GPT-5.6 Sol's 37.5%, pricing unchanged at $5/$25 per 1M tokens (fast tier $10/$50). That is the "we cost more and we're worth it" bet — see our Opus 5 vs Kimi K3 analysis for the compliance subplot on the open-weight side.

App Layer: What Developers Actually Build on OpenRouter

Model rankings show which brain is popular. The apps leaderboard shows what that brain does in production.

RankAppTypeShare (approx.)
1Hermes AgentPersonal / CLI agent (Nous Research)~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent Gateway~9%
4Claude CodeCoding agent (Anthropic)~6%
5DescriptContent production~4.5%
6piAgent~3.3%
7LemonadeCompanion / gaming~2.1%
8ISEKAI ZERORoleplay~2.0%
9Janitor AIRoleplay~1.8%
10ClineIDE coding agent~1.7%

Three takeaways for builders:

  1. Coding agents dominate the long tail after Hermes — Kilo Code, OpenClaw, Claude Code, and Cline all rank in the top ten.
  2. Cline → Roo Code → Kilo Code is one open-source lineage across three forks; the youngest fork (Kilo Code) now beats its ancestors on volume. First-mover advantage in dev tooling is thin.
  3. Roleplay and companion apps (Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI) move serious token volume that enterprise AI reporting almost never covers. OpenRouter's State of AI research with a16z found creative roleplay accounts for more than half of open-model usage on the platform.

Pricing: Cheapest LLM API for Coding vs Premium Tiers

ModelInput / MOutput / MContextPositioning
DeepSeek V4 Flash~$0.05-0.14~$0.24-0.281MCheapest LLM API for coding at scale; daily volume leader
Nemotron 3 Ultra$0.42 (free tier available)$2.61U.S. open-weight; NVIDIA ecosystem
MiniMax M3$0.10$1.21LongBudget multimodal / image input
GLM 5.2$0.45$3.31Closest open-weight match to Opus-style planning
Kimi K3~$3~$151MLargest open weights (1.4TB); fastest July riser
Claude Opus 5$5 ($10 fast)$25 ($50 fast)1MClosed frontier; FrontierBench leader (7/24)

August 2026 Outlook: Five Bets Worth Tracking

  1. Chinese combined share likely crosses 50% unless a major U.S. vendor cuts list pricing. No public signal yet from OpenAI or Google.
  2. The monthly #1 model keeps rotating. Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot ship fast; expect another surprise top-slot contender in August.
  3. Anthropic may ship a cheaper volume tier (Haiku-class) instead of relying on Opus 5 alone. Four flagship releases in under two months (Mythos 5, Fable 5, Sonnet 5, Opus 5) already signal full price-ladder coverage.
  4. Kimi K3's 1.4TB weights should see community quantization within 2-4 weeks, following prior mega-release patterns. Until then, practical access stays with large inference hosts, not local laptops.
  5. Security and governance enter selection scorecards. OpenAI's sandbox-escape incident, the proposed AI Kill Switch Act, and a White House pre-release review framework expected before August 1 make vendor safety history a formal procurement line item — favoring labs with cleaner records like Anthropic for regulated agent workloads.

Six-Step Runbook: Tiered Routing from July Rankings Data

  1. Tag workloads before picking models. Split traffic into bulk, code, agent-long, and critical. One default model for everything is how FinOps teams lose budget.
  2. Map primaries to July volume leaders by tag. Bulk and routine code: DeepSeek V4 Flash or Mimo V2.5. Planning-heavy open steps: GLM 5.2. Critical classification and compliance: Claude Opus 5 or Sonnet 5.
  3. Configure fallback chains in your Gateway. Use the OpenRouter API guide for model IDs and provider routing; never run production without 429/timeout failover.
  4. Isolate free-model queues. Route Nemotron 3 Ultra free and other $0 tiers only to non-sensitive bulk experiments. Block them on production critical paths.
  5. Reconcile weekly against OpenRouter movers. Compare public ranking shifts with your error-rate logs. Cheap tokens with rising failures mean wrong routing, not a bargain.
  6. Probe 24/7 from a stable host. Long Hermes, Kilo Code, or OpenClaw sessions need always-on Gateways. Pin routing to a dedicated node so laptop sleep does not break agent chains — see MACCOME rental rates for M4/M4 Pro tiers.
python
# Tiered OpenRouter routing — July 2026 leaderboard picks
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
)

ROUTES = {
    "bulk": {
        "model": "deepseek/deepseek-v4-flash",
        "fallback": ["xiaomi/mimo-v2.5", "tencent/hy3"],
    },
    "code": {
        "model": "deepseek/deepseek-v4-flash",
        "fallback": ["z-ai/glm-5.2", "anthropic/claude-sonnet-5"],
    },
    "agent-long": {
        "model": "moonshotai/kimi-k3",
        "fallback": ["deepseek/deepseek-v4-pro", "minimax/minimax-m3"],
    },
    "critical": {
        "model": "anthropic/claude-opus-5",
        "fallback": ["google/gemini-3-flash-preview"],
    },
}

def chat(task: str, messages: list[dict]) -> str:
    route = ROUTES[task]
    models = [route["model"], *route["fallback"]]
    resp = client.chat.completions.create(
        model=models[0],
        extra_body={"models": models},  # OpenRouter fallback chain
        messages=messages,
    )
    return resp.choices[0].message.content

Three Hard Numbers for Your Technical Review

  • 46% Chinese share, 12-month climb: China-origin labs rose from under 2% to about 46% of OpenRouter token volume while US labs fell from ~70% to 30-36% — one of the steepest vendor-share migrations in the model API era.
  • 35x price scissors: DeepSeek V4 Flash input at ~$0.05-$0.14/M versus GPT-5.5 at ~$5/M. That gap, not benchmark marketing, explains why Mimo V2.5 and DeepSeek own the daily volume board.
  • 45% single-app concentration: Hermes Agent alone holds roughly 45% of tracked app token share — a reminder that one high-traffic agent stack can distort model rankings more than a thousand indie API calls.

Capability and Popularity Are Diverging — Route Accordingly

July's OpenRouter story is not simply "Chinese models won." It is margin compression at the volume layer while frontier labs defend hard-task pricing and safety credibility. Chinese open models bought half the traffic with price. US closed models still collect premium spend where mistakes are expensive.

For most developers and tech leads, the useful skill is not memorizing today's #1. It is building architecture that swaps models in hours — because the leader on July 25 may not be the leader in August. Pair monthly OpenRouter reviews with your own eval set (acceptance rate, error rate, latency under real load), and encode the barbell in tiered routing instead of picking one "best LLM July 2026" for every task.

Running Hermes, Kilo Code, OpenClaw, or custom Gateways on a laptop hides three costs: lid-close sleep, network jitter, and scattered API keys. Production agent stacks need a dedicated always-on node. MACCOME Mac cloud hosts provide real macOS, SSH handoff, and environment isolation so multi-provider OpenRouter routing stays stable 24/7. Plans and pricing: Mac mini cloud rental rates.

Sources: OpenRouter official rankings and apps leaderboard, OpenRouter State of AI (a16z), Anthropic Opus 5 release, third-party mirrors (tokenmaxxing, presenc.ai). All figures as of July 25, 2026 unless noted.

FAQ

What is the most popular model on OpenRouter in July 2026?

By daily token volume as of July 25: Xiaomi Mimo V2.5 (1.4T/day), then DeepSeek V4 Flash (943.9B) and Tencent Hy3 (590B). By weekly vendor share, DeepSeek remains the most stable #1 at roughly 16-18%. Live data: openrouter.ai/rankings.

Are OpenRouter rankings explained as a quality leaderboard?

No. They rank paid token throughput — developers voting with wallets. For best LLM July 2026 on hard tasks, look at spend on classification and reasoning (Claude Sonnet 4.6, Opus 4.7, GPT-5.5) and vendor benchmarks like FrontierBench, not daily volume alone. Our June rankings deep-dive covers the same volume-vs-quality split with H2 context.

What is the cheapest LLM API for coding right now?

DeepSeek V4 Flash at roughly $0.05-$0.14 input and $0.24-$0.28 output per million tokens, with 1M context. It leads July daily volume and fits bulk agentic coding. Use GLM 5.2 when you need stronger open-weight planning; escalate to Claude Opus 5 only on steps cheaper models fail.

Does 46% Chinese AI model market share mean replace all US models?

No. Route by task tier: Chinese open models for bulk chat, roleplay, and routine code; US frontier models for classification, compliance, and high-stakes agent decisions. Token share and dollar spend tell different stories — Anthropic still captures premium spend on hard workloads even at lower volume.

How do I run tiered OpenRouter routing 24/7 in production?

Deploy a Gateway with task-tagged primaries and fallbacks on a dedicated Mac cloud host so long agent sessions survive sleep and network blips. See MACCOME Mac cloud rental plans for M4/M4 Pro node configs and monthly pricing.

What changed since the June OpenRouter rankings?

Mimo V2.5 overtook DeepSeek V4 Flash for daily #1; Kimi K3 entered the top 12 with the fastest 30-day rise; Claude Opus 5 shipped July 24 as the new premium anchor. Compare month-over-month in our June 2026 rankings article.