OPENROUTER
JULY_2026_
XIAOMI_TOP1_
CHINA_46%.

OpenRouter July 2026 rankings chart showing Chinese model market share

If you are still picking an LLM from a benchmark chart you saw two months ago, you are already behind. Pain point: OpenRouter ranks real paid token volume, not capability scores — and #1 on the volume chart is not #1 on quality. Conclusion: As of July 25, Chinese labs hold roughly 46% of identified token share; Xiaomi's Mimo V2.5 leads at 1.4T tokens/day, and the market is splitting into a barbell of cheap volume vs. premium hard tasks. Structure: Top 12 table → provider share → usage vs. quality → app leaderboard → pricing matrix → August outlook → 5-step routing → Mac three-tier split.

1. Pain Points: Why July's Data Should Rewire Your Routing Stack

1) Volume ≠ capability — a cheap model wired into one high-traffic app can outrank a genuinely stronger model reserved for the hardest 10% of work. 2) The "model of the month" keeps rotating — DeepSeek is the most stable #1 provider (16–18% share), but the daily model crown has moved from MiniMax M2.5 to MiMo-V2-Pro to Mimo V2.5 in July alone. 3) Half the market is invisible in enterprise coverage — roleplay and companion apps move serious open-model volume. 4) Pricing gap ≈ 35× — DeepSeek V4 Flash at ~$0.05–0.14/M input vs. GPT-5.5 at ~$5/M. 5) Security is now a selection variable — after OpenAI's sandbox escape incident, vendor safety track records belong on enterprise scorecards.

2. Model Token Volume Top 12 (Through July 25, 2026)

RankModelLabDaily Tokens30-Day Total
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra 550B (free)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot AI157.6B1.6T (new entry)
10Ling 3.0 FlashInclusionAI128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

Seven of the top ten model slots belong to Chinese labs. Kimi K3 is the fastest riser — the key variable to watch heading into August.

3. Provider Share: China at ~46%, Up From Under 2% a Year Ago

ProviderOriginToken Share (approx.)
DeepSeekChina16%–18% (#1 in most windows)
XiaomiChina8%–18% (most volatile — Mimo V2.5 spike)
AnthropicUS10%–15%
TencentChina8%–13%
GoogleUS8%–13%
Z.aiChina4%–7%
OpenAIUS6%–8%

Chinese labs combined: ~46% (under 2% a year ago). US frontier trio combined: ~30%–36% (down from ~70% in mid-2025). This is one of the steepest share migrations in AI history.

4. What Most Coverage Misses: Usage Rank Is Not a Quality Signal

By spend category: general chat 35.7%, agentic workflows 30.4%, code 26.5%, data work 7.5%. Drill into the hardest bucket — classification/complex reasoning — and the picture flips: Claude Sonnet 4.6 and Claude Opus 4.7 tie at 13.5% of spend each; GPT-5.5 is third at 11.6%. The cheap open models dominating volume charts barely register here.

On July 24, Anthropic shipped Claude Opus 5: 43.3% on FrontierBench v0.1 (vs. GPT-5.6 Sol's 37.5%), holding Opus-tier pricing at $5/$25 per M — half of Fable 5's input price. A clear "we cost more and we're worth it" bet.

ModelInput/MOutput/MPositioning
DeepSeek V4 Flash~$0.05–0.14~$0.24–0.28Best value; agentic coding default
MiniMax M3$0.10$1.21Long-context multimodal on a budget
GLM 5.2$0.45$3.31Closest open-weight Opus-style planner
Kimi K3~$3~$151.4TB open weights, closed-tier capability
Claude Opus 5$5 (fast tier $10)$25 (fast tier $50)Closed frontier; July benchmark leader

5. App Layer: Coding Agents Dominate; Roleplay Is the Invisible Half

RankAppCategoryShare (approx.)
1Hermes AgentPersonal agent / CLI~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent~9%
4Claude CodeCoding agent~6%
5DescriptContent production~4.5%
7LemonadeCompanion / gaming~2.1% (new)
8ISEKAI ZERORoleplay~2.0% (new)
10ClineCoding agent (IDE)~1.7%

Cline → Roo Code → Kilo Code is three generations of the same open-source lineage — the youngest fork now leads volume. OpenRouter × a16z's State of AI report found creative roleplay accounts for more than half of all open-model usage. If your view of "AI usage" comes only from enterprise headlines, you are missing half the market.

6. August Outlook: Five Signals to Watch

  1. Chinese open-weight combined share likely climbs toward 50% unless a major US provider makes a real pricing move.
  2. The monthly #1 model keeps rotating — Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot are all shipping and re-pricing fast.
  3. Anthropic may ship a cheaper volume tier — Opus 5 is their fourth flagship in under two months (after Mythos 5, Fable 5, Sonnet 5), signaling full price-tier coverage.
  4. Kimi K3's 1.4TB weights will likely see community quantization within 2–4 weeks — until then, practical access is mostly for large teams and inference hosts like Fireworks AI.
  5. Security and governance become formal selection criteria — OpenAI's sandbox escape, the proposed AI Kill Switch Act, and a White House pre-release review framework expected before August 1.

7. Five Steps: Build Task-Tier Routing That Survives August

  1. Map chains by task: chat/creative/roleplay → DeepSeek V4 Flash / Mimo V2.5; classification/complex reasoning/high-stakes agents → Claude Opus 5 / GPT-5.6.
  2. Set spend caps: daily token budgets on frontier models; auto-fallback to GLM 5.2 or DeepSeek Flash on breach.
  3. Check openrouter.ai/rankings weekly: rankings shift daily — verify Top 10 before citing; preview pricing often ends with rank spikes.
  4. Local open-weight backup: GLM 5.2 / Kimi K3 on Mac via MLX for offline fallback during rate limits or compliance switches.
  5. Run a 20-task eval set: same tasks on cheap and frontier models; log pass rate and cost per task into team SOP.

8. Deep Dive: Capability and Popularity Are Diverging

The line to remember from July: capability and popularity are splitting apart. Chinese open-weight models bought half the market with price — a ~35× gap between DeepSeek V4 Flash and GPT-5.5 when 80–90% of daily tasks are "good enough." US closed-frontier labs defend the other half with pricing power on hard tasks: Claude Opus 5's 43.3% FrontierBench score and Anthropic's relatively clean safety record are why enterprises still pay premium rates.

At the app layer, Hermes Agent holds 45% of tracked app token share; the coding agent family (Kilo Code, OpenClaw, Claude Code, Cline) fills most of the rest. For Mac developers, routing tables must split by workload — Cursor completions, OpenClaw agent chains, and Hermes self-improvement loops often need different optimal model IDs.

Independent developers: OpenRouter remains the fastest way to A/B test dozens of models behind one API key. Start coding workloads with DeepSeek V4 Flash for cost and GLM 5.2 for Opus-style planning; reserve premium closed models for steps where cheaper models actually fail.

9. Close: Tiered Routing + Mac Unified Memory

A Windows/Linux cloud box can call OpenRouter APIs, but it falls short on local MLX inference, Cursor toolchain synergy, 24/7 agent uptime, and graphics workflows compared to Apple Silicon Mac. If you are navigating the "DeepSeek cents vs. Claude dollars per hour" gap and need predictable local backup plus remote peak offload, use a three-tier stack: local MLX for GLM 5.2 / Kimi K3 daily volume; OpenRouter API for Opus 5 on hard steps and DeepSeek Flash on volume; MACGPU remote Mac nodes for overnight batch agents and long-context jobs that stress unified memory — controlled compute is the best hedge before August's model wave.