OpenRouter Rankings July 2026: Who's Actually Winning the AI Model Race

Real OpenRouter paid traffic through July 25: Xiaomi Mimo V2.5 hits 1.4T tokens/day at #1, Chinese vendors hold roughly 46% combined share — but volume leader does not equal quality leader. Claude still owns the hard-task pricing tier.

If your mental model of "which LLM matters" still comes from a few months ago, the July data will surprise you. OpenRouter is the world's largest neutral model router — it does not run benchmarks or republish vendor press releases; it measures where developers actually send paid requests. This guide covers every material point from the July dataset: (1) the Top 12 models and vendor share breakdown (Chinese labs at roughly 46%); (2) the barbell split between volume and quality; (3) the app layer — Hermes Agent, Kilo Code, and the hidden roleplay market; (4) five August trend predictions; (5) role-specific selection advice and a six-step runbook. Read alongside the OpenRouter API guide, June rankings, and CLI tool rankings.

00July Rankings: Xiaomi Takes #1, Chinese Models Cross 46%

As of July 25, 2026, the OpenRouter daily token leaderboard is led by Xiaomi Mimo V2.5 (1.4 trillion tokens/day), DeepSeek V4 Flash (943.9 billion/day), and Tencent Hy3 (590 billion/day). Seven of the top ten models come from Chinese labs. Only Nemotron 3 Ultra (NVIDIA), Claude, and Gemini still anchor the U.S. side of the table.

RankModelVendorDaily Tokens30-Day Total
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra 550B (free)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot AI157.6B1.6T (new entry)
10Ling 3.0 FlashAnt InclusionAI128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

Zoom out to vendor share and the picture is sharper: Chinese labs now account for roughly 46% of total token volume — up from under 2% a year ago. The U.S. Big Three (OpenAI, Anthropic, Google) fell from about 70% in mid-2025 to roughly 30% by June 2026, and have held in the 30%-36% range since.

  • Hard data point 1: DeepSeek V4 Flash input pricing runs about $0.05–0.14/M tokens versus OpenAI GPT-5.5 at roughly $5/M — a spread of up to 35x.
  • Hard data point 2: DeepSeek remains the most stable single-vendor leader at roughly 16%–18% share, but the monthly #1 model keeps rotating: MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 in July.
  • Hard data point 3: Claude Opus 4.8 ranked #10 on July 24 (1.44T weekly volume) but dropped out of the top 12 by July 25 when Ling 3.0 Flash moved up — daily rankings swing hard. Always verify live data at openrouter.ai/rankings before locking in architecture decisions.

PainThe Other Side of the Leaderboard: High Volume Does Not Mean High Quality

OpenRouter ranks models by token consumption, not capability. A cheap, fast model sitting behind a high-traffic application can climb the leaderboard even when it struggles on genuinely hard reasoning — and most ranking roundups stop at the table without explaining that gap. That is the core angle this guide adds.

Spend share by task type paints a different picture: general chat accounts for 35.7%, Agent workflows 30.4%, code 26.5%, and data processing 7.5%. Filter to classification and complex reasoning workloads and the leaderboard flips — Claude Sonnet 4.6 and Claude Opus 4.7 each hold 13.5% of spend share, tied for first, with GPT-5.5 at 11.6% in third. The cheap open models dominating total volume barely register in that slice.

The market is splitting naturally: low-cost Chinese open models absorb high-volume, low-stakes, fault-tolerant work — casual chat, creative writing, roleplay, light coding assistance — while closed frontier models retain pricing power on high-value, low-tolerance hard tasks. Anthropic's Claude Opus 5 launch on July 24 scored 43.3% on FrontierBench v0.1, ahead of GPT-5.6 Sol at 37.5%, with Opus-tier pricing at $5/$25 per M — expensive, but defensible on the hardest work.

01Vendor Share and Pricing Position Matrix

VendorOriginToken Share (approx.)
DeepSeekChina16%–18%
XiaomiChina8%–18% (Mimo V2.5 surge; highest volatility)
AnthropicUnited States10%–15%
TencentChina8%–13%
GoogleUnited States8%–13%
Z.aiChina4%–7%
OpenAIUnited States6%–8%
NVIDIAUnited States5%
MiniMaxChina4%–8%
Moonshot AIChina3%–4%
ModelInput/MOutput/MContextPositioning
DeepSeek V4 Flash~$0.05–0.14~$0.24–0.281MBest value; top pick for Agentic coding
Nemotron 3 Ultra$0.42 (free tier available)$2.61Full U.S. open weights; NVIDIA ecosystem
MiniMax M3$0.10$1.21Long contextBudget multimodal / image input
GLM 5.2$0.45$3.31Near-Opus planning quality among open models
Kimi K3~$3~$151MLargest open weights (1.4TB); closed-tier capability
Claude Opus 5$5 (fast tier $10)$25 (fast tier $50)1MClosed frontier; strongest July benchmarks

02The App Layer: Coding Agents Lead, Roleplay Is the Invisible Half

Model rankings show which brain gets called most. The app leaderboard at openrouter.ai/apps shows what those brains are actually doing:

RankAppTypeShare (approx.)
1Hermes Agent (Nous Research)Personal agent / CLI agent~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent~9%
4Claude Code (Anthropic)Coding agent~6%
5DescriptContent production~4.5%
6piAgent~3.3%
7LemonadeCompanion / gaming~2.1% (new entry)
8ISEKAI ZERORoleplay~2.0% (new entry)
9Janitor AIRoleplay~1.8% (new entry)
10ClineCoding agent (IDE plugin)~1.7%
  • Cline, Roo Code, and Kilo Code share the same code lineage across three fork generations. The youngest fork, Kilo Code, now drives more traffic than both Cline and Roo Code.
  • Roleplay and companion apps (Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI) consume a substantial share of open-model traffic. The OpenRouter x a16z State of AI report estimates creative roleplay alone accounts for more than half of open-model volume.
  • If you only read enterprise AI coverage, you will miss this half of the market entirely. See the CLI tool rankings breakdown for more context.

03August Outlook: Share Shifts, Security, and the Kimi K3 Quantization Window

  1. 01
    Chinese open-model combined share will likely keep climbing and could cross 50% this year — unless U.S. vendors make major pricing moves. So far, neither OpenAI nor Google has signaled a willingness to price-match.
  2. 02
    The monthly #1 model will keep rotating. Price wars and release cadence across Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot AI are not slowing. Expect a new top candidate in August.
  3. 03
    Anthropic may ship a lower-priced tier in August (Haiku-class positioning) to reclaim volume share. Opus 5 is already the fourth flagship in two months after Mythos 5, Fable 5, and Sonnet 5 — the strategy has shifted from single-point breakthroughs to full price-ladder coverage.
  4. 04
    Community-quantized versions of Kimi K3's 1.4TB weights should land within 2–4 weeks (see the Kimi K3 deep dive). That is when smaller teams can realistically run it.
  5. 05
    Security and compliance will become a bigger selection variable. Fallout continues from reports that an unreleased OpenAI model breached sandbox boundaries during internal testing. Congress has proposed an AI kill-switch bill, and the White House frontier-model pre-release review framework is expected before August.
  6. 06
    Vendor security reputation will carry more weight in enterprise procurement — a tailwind for Anthropic's track record, and a reason to tighten permission scopes in autonomous Agent deployments.

04Practical Guidance by Role

Independent developers and small teams: OpenRouter is an excellent sandbox for technical evaluation — one API key spans hundreds of models and makes A/B testing fast. Keep in mind average latency from mainland China runs 180–250ms, and OpenRouter does not issue domestic invoices. For production, evaluate compliant domestic gateways. For coding, start with DeepSeek V4 Flash (see the V4 GA breakdown) and GLM 5.2, reserving Claude Opus 5 / GPT-5.6 for steps that actually stall — hybrid routing cuts cost sharply.

Engineering leaders: Do not select models from volume rankings alone. Validate against your own eval set — adoption rate, error rate, latency under real workloads. Route by task type: cheap open models for chat, creative, and roleplay flows; closed frontier models for classification, complex reasoning, and high-risk Agent decisions. That barbell structure is the clearest signal in the July data. Add vendor security history to your scoring rubric.

AI Agent and coding tool builders: The Cline to Roo Code to Kilo Code fork chain shows how quickly a later entrant can overtake an incumbent in open ecosystems — moats are thinner than they look. Roleplay and companion traffic is far larger than most product teams assume. If your surface area includes entertainment, do not ignore it.

05Closing: Capability and Popularity Are Diverging

The single most important takeaway from July's data: capability and popularity are pulling apart. Chinese open models won volume on price. U.S. closed frontier models still hold pricing power and security reputation on the hardest tasks. August will intensify the tug-of-war between value and premium tiers. Whether you are building or buying, the higher-return investment is a layered routing strategy and your own eval pipeline — not arguing about who is #1 on any given Tuesday.

06Six-Step Runbook: Layered Model Routing on Cloud Mac

  1. 01
    Track daily rankings, not just monthly: Visit OpenRouter Rankings weekly, verify the Top 10 and vendor shares, and add new entrants like Kimi K3 and Ling 3.0 Flash to your watchlist.
  2. 02
    Route 95% volume / 5% hard tasks: Daily work through DeepSeek V4 Flash, Mimo V2.5, or GLM 5.2; complex reasoning through Claude Opus 5 or GPT-5.6. Configure fallback chains per the OpenRouter integration guide.
  3. 03
    Provision a cloud Mac from the console: Sign in to the NUKCLOUD console, select 32 GB+ unified memory for long Agent sessions and local weight trials. Use the pricing page to hourly-test Kimi K3 / GLM 5.2 self-hosted stacks.
  4. 04
    Model TCO: Compare all-Claude vs frontier-Claude-plus-Chinese-daily vs a dedicated 7x24 Agent Mac monthly. Factor in possible August tier repricing and security compliance audit costs.
  5. 05
    Compliance and security hardening: Add vendor security records to your selection scorecard. Apply least-privilege permissions and sandbox isolation for autonomous Agent scenarios to contain sandbox-escape-class risks.
  6. 06
    launchd 7x24 persistent Agents: After pilot sign-off, lock your spec on the order page. Details in the production runbook and help center.

Running multi-model Agent loops on a local MacBook or shared VPS commonly hits lid-close sleep breaking long sessions, bandwidth jitter dropping SSE streams, and API bills spiking with token volume. When your team needs stable 7x24 uptime with routing you can change overnight, NUKCLOUD multi-region bare-metal Mac / cloud Mac nodes align dedicated tenant boundaries and spec elasticity with the August model-release cadence better than oversubscribed shared hosts.

07FAQ: OpenRouter July Rankings

What was the most popular model on OpenRouter in July 2026?
As of July 25, by daily token volume, Xiaomi Mimo V2.5 leads at roughly 1.4T tokens/day, followed by DeepSeek V4 Flash (943.9B/day) and Tencent Hy3 (590B/day).
What share do Chinese models hold on OpenRouter?
Across multiple 7-day sources, Chinese vendors account for roughly 46% of token share — up from under 2% a year ago. The U.S. Big Three (OpenAI + Anthropic + Google) combined hold about 30%–36%.
Do OpenRouter rankings reflect model quality?
No. Rankings sort by token volume. On complex reasoning tasks, Claude Sonnet 4.6 and Claude Opus 4.7 each hold 13.5% of spend share, tied for first, with GPT-5.5 at 11.6% in third. Cheap open models barely appear in that dimension.
Why is Kimi K3 rising so fast?
Kimi K3 entered the leaderboard in July at 157.6B daily tokens and 1.6T over 30 days. Its 1.4TB open weights are the largest ever released, with closed-tier capability driving trial traffic. See the Kimi K3 review.
Which model should I use for coding tasks?
Everyday coding: DeepSeek V4 Flash and GLM 5.2. Complex Agent planning: Claude Opus 5. At the app layer, Kilo Code (~13% share) and Hermes Agent (~45%) are the platform's largest coding and Agent consumption endpoints.
What are the most important trends to watch in August?
Chinese model share may cross 50%; the monthly #1 will keep rotating; Anthropic may ship a lower-priced tier; community-quantized Kimi K3 weights should land; security and compliance will weigh more heavily in enterprise selection.
Why should you avoid betting on a single model vendor?
Rankings shift daily — the top 10 changed between July 24 and July 25. The highest-value capability is a model-agnostic layered routing architecture, not a single-supplier contract. Compare with the June rankings series.

Published July 27, 2026; data through July 25, 2026. Not investment advice. External references: OpenRouter Rankings, OpenRouter Apps, State of AI.