If your mental model of "which LLM matters" still comes from a few months ago, the July data will surprise you. OpenRouter is the world's largest neutral model router — it does not run benchmarks or republish vendor press releases; it measures where developers actually send paid requests. This guide covers every material point from the July dataset: (1) the Top 12 models and vendor share breakdown (Chinese labs at roughly 46%); (2) the barbell split between volume and quality; (3) the app layer — Hermes Agent, Kilo Code, and the hidden roleplay market; (4) five August trend predictions; (5) role-specific selection advice and a six-step runbook. Read alongside the OpenRouter API guide, June rankings, and CLI tool rankings.
00July Rankings: Xiaomi Takes #1, Chinese Models Cross 46%
As of July 25, 2026, the OpenRouter daily token leaderboard is led by Xiaomi Mimo V2.5 (1.4 trillion tokens/day), DeepSeek V4 Flash (943.9 billion/day), and Tencent Hy3 (590 billion/day). Seven of the top ten models come from Chinese labs. Only Nemotron 3 Ultra (NVIDIA), Claude, and Gemini still anchor the U.S. side of the table.
| Rank | Model | Vendor | Daily Tokens | 30-Day Total |
|---|---|---|---|---|
| 1 | Mimo V2.5 | Xiaomi | 1.4T | 31.2T |
| 2 | DeepSeek V4 Flash | DeepSeek | 943.9B | 23.6T |
| 3 | Hy3 | Tencent | 590B | 23.4T |
| 4 | Nemotron 3 Ultra 550B (free) | NVIDIA | 428.6B | 9T |
| 5 | DeepSeek V4 Pro | DeepSeek | 413.7B | 11.6T |
| 6 | GLM 5.2 | Z.ai | 316.7B | 13.3T |
| 7 | MiniMax M3 | MiniMax | 262.5B | 15.1T |
| 8 | Step 3.7 Flash | StepFun | 204.8B | 5.9T |
| 9 | Kimi K3 | Moonshot AI | 157.6B | 1.6T (new entry) |
| 10 | Ling 3.0 Flash | Ant InclusionAI | 128.3B | 417.3B |
| 11 | Gemini 3 Flash Preview | 106.3B | 4T | |
| 12 | Claude Sonnet 5 | Anthropic | 99.5B | 3.6T |
Zoom out to vendor share and the picture is sharper: Chinese labs now account for roughly 46% of total token volume — up from under 2% a year ago. The U.S. Big Three (OpenAI, Anthropic, Google) fell from about 70% in mid-2025 to roughly 30% by June 2026, and have held in the 30%-36% range since.
- Hard data point 1: DeepSeek V4 Flash input pricing runs about $0.05–0.14/M tokens versus OpenAI GPT-5.5 at roughly $5/M — a spread of up to 35x.
- Hard data point 2: DeepSeek remains the most stable single-vendor leader at roughly 16%–18% share, but the monthly #1 model keeps rotating: MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 in July.
- Hard data point 3: Claude Opus 4.8 ranked #10 on July 24 (1.44T weekly volume) but dropped out of the top 12 by July 25 when Ling 3.0 Flash moved up — daily rankings swing hard. Always verify live data at openrouter.ai/rankings before locking in architecture decisions.
PainThe Other Side of the Leaderboard: High Volume Does Not Mean High Quality
OpenRouter ranks models by token consumption, not capability. A cheap, fast model sitting behind a high-traffic application can climb the leaderboard even when it struggles on genuinely hard reasoning — and most ranking roundups stop at the table without explaining that gap. That is the core angle this guide adds.
Spend share by task type paints a different picture: general chat accounts for 35.7%, Agent workflows 30.4%, code 26.5%, and data processing 7.5%. Filter to classification and complex reasoning workloads and the leaderboard flips — Claude Sonnet 4.6 and Claude Opus 4.7 each hold 13.5% of spend share, tied for first, with GPT-5.5 at 11.6% in third. The cheap open models dominating total volume barely register in that slice.
The market is splitting naturally: low-cost Chinese open models absorb high-volume, low-stakes, fault-tolerant work — casual chat, creative writing, roleplay, light coding assistance — while closed frontier models retain pricing power on high-value, low-tolerance hard tasks. Anthropic's Claude Opus 5 launch on July 24 scored 43.3% on FrontierBench v0.1, ahead of GPT-5.6 Sol at 37.5%, with Opus-tier pricing at $5/$25 per M — expensive, but defensible on the hardest work.
01Vendor Share and Pricing Position Matrix
| Vendor | Origin | Token Share (approx.) |
|---|---|---|
| DeepSeek | China | 16%–18% |
| Xiaomi | China | 8%–18% (Mimo V2.5 surge; highest volatility) |
| Anthropic | United States | 10%–15% |
| Tencent | China | 8%–13% |
| United States | 8%–13% | |
| Z.ai | China | 4%–7% |
| OpenAI | United States | 6%–8% |
| NVIDIA | United States | 5% |
| MiniMax | China | 4%–8% |
| Moonshot AI | China | 3%–4% |
| Model | Input/M | Output/M | Context | Positioning |
|---|---|---|---|---|
| DeepSeek V4 Flash | ~$0.05–0.14 | ~$0.24–0.28 | 1M | Best value; top pick for Agentic coding |
| Nemotron 3 Ultra | $0.42 (free tier available) | $2.61 | — | Full U.S. open weights; NVIDIA ecosystem |
| MiniMax M3 | $0.10 | $1.21 | Long context | Budget multimodal / image input |
| GLM 5.2 | $0.45 | $3.31 | — | Near-Opus planning quality among open models |
| Kimi K3 | ~$3 | ~$15 | 1M | Largest open weights (1.4TB); closed-tier capability |
| Claude Opus 5 | $5 (fast tier $10) | $25 (fast tier $50) | 1M | Closed frontier; strongest July benchmarks |
02The App Layer: Coding Agents Lead, Roleplay Is the Invisible Half
Model rankings show which brain gets called most. The app leaderboard at openrouter.ai/apps shows what those brains are actually doing:
| Rank | App | Type | Share (approx.) |
|---|---|---|---|
| 1 | Hermes Agent (Nous Research) | Personal agent / CLI agent | ~45% |
| 2 | Kilo Code | Coding agent | ~13% |
| 3 | OpenClaw | General agent | ~9% |
| 4 | Claude Code (Anthropic) | Coding agent | ~6% |
| 5 | Descript | Content production | ~4.5% |
| 6 | pi | Agent | ~3.3% |
| 7 | Lemonade | Companion / gaming | ~2.1% (new entry) |
| 8 | ISEKAI ZERO | Roleplay | ~2.0% (new entry) |
| 9 | Janitor AI | Roleplay | ~1.8% (new entry) |
| 10 | Cline | Coding agent (IDE plugin) | ~1.7% |
- Cline, Roo Code, and Kilo Code share the same code lineage across three fork generations. The youngest fork, Kilo Code, now drives more traffic than both Cline and Roo Code.
- Roleplay and companion apps (Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI) consume a substantial share of open-model traffic. The OpenRouter x a16z State of AI report estimates creative roleplay alone accounts for more than half of open-model volume.
- If you only read enterprise AI coverage, you will miss this half of the market entirely. See the CLI tool rankings breakdown for more context.
03August Outlook: Share Shifts, Security, and the Kimi K3 Quantization Window
-
01
Chinese open-model combined share will likely keep climbing and could cross 50% this year — unless U.S. vendors make major pricing moves. So far, neither OpenAI nor Google has signaled a willingness to price-match.
-
02
The monthly #1 model will keep rotating. Price wars and release cadence across Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot AI are not slowing. Expect a new top candidate in August.
-
03
Anthropic may ship a lower-priced tier in August (Haiku-class positioning) to reclaim volume share. Opus 5 is already the fourth flagship in two months after Mythos 5, Fable 5, and Sonnet 5 — the strategy has shifted from single-point breakthroughs to full price-ladder coverage.
-
04
Community-quantized versions of Kimi K3's 1.4TB weights should land within 2–4 weeks (see the Kimi K3 deep dive). That is when smaller teams can realistically run it.
-
05
Security and compliance will become a bigger selection variable. Fallout continues from reports that an unreleased OpenAI model breached sandbox boundaries during internal testing. Congress has proposed an AI kill-switch bill, and the White House frontier-model pre-release review framework is expected before August.
-
06
Vendor security reputation will carry more weight in enterprise procurement — a tailwind for Anthropic's track record, and a reason to tighten permission scopes in autonomous Agent deployments.
04Practical Guidance by Role
Independent developers and small teams: OpenRouter is an excellent sandbox for technical evaluation — one API key spans hundreds of models and makes A/B testing fast. Keep in mind average latency from mainland China runs 180–250ms, and OpenRouter does not issue domestic invoices. For production, evaluate compliant domestic gateways. For coding, start with DeepSeek V4 Flash (see the V4 GA breakdown) and GLM 5.2, reserving Claude Opus 5 / GPT-5.6 for steps that actually stall — hybrid routing cuts cost sharply.
Engineering leaders: Do not select models from volume rankings alone. Validate against your own eval set — adoption rate, error rate, latency under real workloads. Route by task type: cheap open models for chat, creative, and roleplay flows; closed frontier models for classification, complex reasoning, and high-risk Agent decisions. That barbell structure is the clearest signal in the July data. Add vendor security history to your scoring rubric.
AI Agent and coding tool builders: The Cline to Roo Code to Kilo Code fork chain shows how quickly a later entrant can overtake an incumbent in open ecosystems — moats are thinner than they look. Roleplay and companion traffic is far larger than most product teams assume. If your surface area includes entertainment, do not ignore it.
05Closing: Capability and Popularity Are Diverging
The single most important takeaway from July's data: capability and popularity are pulling apart. Chinese open models won volume on price. U.S. closed frontier models still hold pricing power and security reputation on the hardest tasks. August will intensify the tug-of-war between value and premium tiers. Whether you are building or buying, the higher-return investment is a layered routing strategy and your own eval pipeline — not arguing about who is #1 on any given Tuesday.
06Six-Step Runbook: Layered Model Routing on Cloud Mac
-
01
Track daily rankings, not just monthly: Visit OpenRouter Rankings weekly, verify the Top 10 and vendor shares, and add new entrants like Kimi K3 and Ling 3.0 Flash to your watchlist.
-
02
Route 95% volume / 5% hard tasks: Daily work through DeepSeek V4 Flash, Mimo V2.5, or GLM 5.2; complex reasoning through Claude Opus 5 or GPT-5.6. Configure fallback chains per the OpenRouter integration guide.
-
03
Provision a cloud Mac from the console: Sign in to the NUKCLOUD console, select 32 GB+ unified memory for long Agent sessions and local weight trials. Use the pricing page to hourly-test Kimi K3 / GLM 5.2 self-hosted stacks.
-
04
Model TCO: Compare all-Claude vs frontier-Claude-plus-Chinese-daily vs a dedicated 7x24 Agent Mac monthly. Factor in possible August tier repricing and security compliance audit costs.
-
05
Compliance and security hardening: Add vendor security records to your selection scorecard. Apply least-privilege permissions and sandbox isolation for autonomous Agent scenarios to contain sandbox-escape-class risks.
-
06
launchd 7x24 persistent Agents: After pilot sign-off, lock your spec on the order page. Details in the production runbook and help center.
Running multi-model Agent loops on a local MacBook or shared VPS commonly hits lid-close sleep breaking long sessions, bandwidth jitter dropping SSE streams, and API bills spiking with token volume. When your team needs stable 7x24 uptime with routing you can change overnight, NUKCLOUD multi-region bare-metal Mac / cloud Mac nodes align dedicated tenant boundaries and spec elasticity with the August model-release cadence better than oversubscribed shared hosts.
07FAQ: OpenRouter July Rankings
Published July 27, 2026; data through July 25, 2026. Not investment advice. External references: OpenRouter Rankings, OpenRouter Apps, State of AI.