Model Comparison

Qwen3.8 vs Kimi K3

China's two trillion-parameter AI models go head-to-head. One has published benchmarks. The other has a bold claim. Here's what the data actually says.

By ZVHH Research July 20, 2026 10 min read

⚡ TL;DR

  • 🚀 Kimi K3 (Moonshot AI, July 16) has published benchmarks across 12 categories, 2.8T parameters, 16/896 MoE, and $3/$15 API pricing
  • 📊 Qwen3.8-Max-Preview (Alibaba, July 19) claims 2.4T parameters and "second only to Fable 5" — but has zero published benchmarks
  • 💰 Qwen's pricing edge is its biggest known advantage: $1.25 input / $3.75 output (Qwen3.7 baseline) vs K3's $3 / $15
  • 🔓 Both promise open weights. K3's deadline is July 27. Qwen3.8's is "soon" — unconfirmed
  • 📈 Kimi K3 leads Qwen3.7-Max on BenchAlign by 8.1 points (80.96 vs 72.84), with a 19.8-point gap in agentic tasks
  • ⚠️ Direct Qwen3.8 vs K3 comparison is currently impossible — Qwen3.8 benchmarks don't exist yet
2.8T
Kimi K3 Parameters
2.4T
Qwen3.8 Parameters (Claimed)
16/896
K3 MoE Sparsity
1M+
Context Window (Both)

The Contenders

Kimi K3 — Moonshot AI

Released July 16, 2026, Kimi K3 is Moonshot AI's flagship open frontier model. At 2.8 trillion total parameters with only 16 of 896 experts active per token, it uses extreme sparsity (less than 2% active) to deliver frontier performance. K3 is available now via the Kimi app, Kimi Code CLI, API, and OpenRouter. Moonshot promises open weights by July 27, 2026. Verified: Moonshot official blog

Qwen3.8-Max-Preview — Alibaba

Announced July 19, 2026, Qwen3.8-Max-Preview is Alibaba's latest Qwen flagship. The Qwen team claims 2.4 trillion parameters and says the model is "second only to Fable 5" among frontier models. It's available through Alibaba's Token Plan ($6/$18/$68/month), Qoder, and QoderWork. Alibaba promises open weights "soon" but has published nothing yet. Unverified: Alibaba claims only

🎯 Why this matters right now

Both models launched within 3 days of each other in July 2026. Both are Chinese labs pushing past the trillion-parameter threshold. Both promise open weights. Both are positioned as alternatives to Claude Fable 5 and GPT 5.6 Sol. The timing is no coincidence — this is a direct competitive confrontation.

Architecture: Published vs Speculated

Kimi K3 Architecture

Moonshot published an extensive architecture blog for K3, detailing several novel components:

Qwen3.8 Architecture

Alibaba has published no architecture details for Qwen3.8. Based on the established Qwen lineage, we can infer:

⚠️ Transparency gap

Kimi K3's architecture is fully documented in Moonshot's blog. Qwen3.8's architecture is entirely speculative. This asymmetry extends to every aspect of the comparison — K3 publishes, Qwen3.8 claims.

Benchmarks: Published vs Claimed

This is where the comparison gets interesting. Kimi K3 has published benchmark tables across 12 categories. Qwen3.8 has published zero benchmarks. What follows is the best available comparison: K3's published data vs Qwen3.7-Max (the previous generation, verified), with Qwen3.8's position estimated based on lineage.

Reasoning & Knowledge

Both models excel here, but K3 has the edge on reasoning benchmarks. The gap narrows on pure knowledge tasks where Qwen3.7-Max remains competitive.

Benchmark K3 (max) Qwen3.7-Max Qwen3.8 Claim
GPQA-Diamond 93.5 92.4
HLE (w/ tools) 56.0 53.5
Artificial Analysis Intelligence Index 57.1 46.0

Coding & Agentic Tasks

K3's agentic coding benchmarks are strong, but Qwen3.7-Max has SWE-bench Verified data that K3 hasn't published on.

Benchmark K3 (max) Qwen3.7-Max Qwen3.8 Claim
Terminal-Bench 2.1 88.3 69.7
SWE-bench Verified 80.4
FrontierSWE 81.2
SWE Marathon 42.0

Agentic & Knowledge Work

K3's biggest advantage is in agentic tasks — a 19.8-point gap over Qwen3.7-Max on BenchAlign's Agentic category.

Benchmark K3 (max) Qwen3.7-Max Qwen3.8 Claim
BrowseComp 91.2
Toolathlon-Verified 73.2
AA Briefcase (Elo) 1548 908
GDPval-AA v2 (Elo) 1668 1273

Multimodal & Vision

Qwen3.8 claims to be the first Qwen model over 1T parameters with multimodal support (images, videos, documents). K3 has native vision with strong MMMU-Pro and CharXiv scores.

Benchmark K3 (max) Qwen3.7-Max Qwen3.8 Claim
MMMU-Pro 81.6 Claimed
CharXiv 84.8
MathVision 94.3

📊 BenchLM comparison (K3 vs Qwen3.7-Max)

BenchAlign aggregate: K3 80.96 vs Q3.7-Max 72.84 (8.1 point margin). K3 leads on Agentic (89.5 vs 69.7) and Knowledge (61.0 vs 64.2 — Qwen3.7 edge). Qwen3.7 leads on Coding, Reasoning, Math, and Multilingual — categories where K3 has no published data yet. Qwen3.8's position vs K3 cannot be measured until benchmarks are published.

Key Benchmark Gaps (K3 vs Qwen3.7-Max)

GPQA-DiamondK3 93.5 vs Q3.7 92.4 (Δ 1.1)
Terminal-Bench 2.0K3 88.3 vs Q3.7 69.7 (Δ 18.6)
AA Briefcase (Elo)K3 1548 vs Q3.7 908 (Δ 640)
SWE-bench VerifiedQ3.7 80.4 vs K3 — (not tested)

Pricing

Pricing is Qwen's strongest known advantage. Even the previous generation (Qwen3.7-Max) is dramatically cheaper than K3. Qwen3.8's pricing is not yet published, but Alibaba has consistently held its value position across generations.

💰 Pricing Comparison

$1.25
Qwen3.7 Input / MTok
$3.75
Qwen3.7 Output / MTok
$3.00
K3 Input / MTok
$15.00
K3 Output / MTok
Model Input / MTok Output / MTok Cache Hit Pricing Tier
Qwen3.7-Max $1.25 $3.75 Value
Kimi K3 $3.00 $15.00 $0.30 Premium
Claude Fable 5 (Sonnet) $10.00 $50.00 Flagship

✅ Qwen's pricing story

Qwen3.7-Max delivers a 92.4 GPQA score at $1.25 input — one eighth of Claude Fable 5's input price and less than half of Kimi K3's. If Qwen3.8 holds this pricing, it will be one of the best value models on the market regardless of where it ranks on benchmarks.

Open Weight Status

Model Status Details
Kimi K3 Promised: July 27, 2026 Deadline set. Community watching. Moonshot has a track record of hitting dates.
Qwen3.8-Max-Preview Promised: "Soon" No date, no license, no HF repo. Would break pattern — last 2 Max flagships were closed.

⚠️ Open-weight credibility check

Both Kimi K3 and Qwen3.8 promise open weights, but K3 has a specific date (July 27) while Qwen3.8 says only "soon." Critically, Alibaba's last two Max-tier flagships (Qwen3.7-Max and Qwen3.6-Max-Preview) both shipped closed through Alibaba Cloud Model Studio. An open-weight Qwen3.8 would break a clear pattern. Treat any promised open-weight release as unconfirmed until the HuggingFace repo exists with a real license file.

Scorecard

Head-to-Head Assessment
Kimi K3
Benchmarks Published
Qwen3.8
Parameter Count (Claimed)
Qwen3.8
Price-to-Performance
Kimi K3
Open Weight Specificity
Kimi K3
Documentation Depth
Qwen3.8
Multimodal (Claimed)

Confidence Ratings

K3 — Benchmarks ★★★★★ 5/10 (self-reported, but extensive and cross-verified on BenchLM/Artificial Analysis)
Qwen3.8 — Claims 1/10 (zero benchmarks published; all performance claims unverified)
Qwen3.8 — Value Prop ★★★★ 8/10 (Qwen3.7 baseline pricing is exceptional; 3.8 likely similar)

Verdict

Current State: Kimi K3 Wins on Evidence, Qwen3.8 on Potential

As of July 20, 2026, Kimi K3 is the more complete product: published benchmarks across 12 categories, documented architecture, accessible API, and a specific open-weight deadline. Qwen3.8-Max-Preview is a promise wrapped in a teaser — the product exists (you can buy access), but every performance claim rests on Alibaba's word alone.

The Qwen3.8 team's claim of "second only to Fable 5" is not verifiable today. It may be true on some benchmarks — Qwen3.7-Max already scores 92.4 on GPQA-Diamond, within 1.1 points of K3's 93.5 — but closing a 15-point gap on SWE-bench Verified (Fable 5 at 95% vs Qwen3.7 at 80.4%) in one generation would be extraordinary.

Qwen3.8's real weapon is pricing. If it delivers even marginally better benchmarks than Qwen3.7-Max while holding the $1.25/$3.75 pricing tier, it will be the best value model on the market regardless of where it ranks on leaderboards.

Who Should Use Which?

✅ Choose Kimi K3 for

Maximum capability today. If you need proven performance on coding, reasoning, agentic tasks, and multimodal work — and you're willing to pay premium pricing — K3 is the stronger model right now. Its open-weight release on July 27 will further expand its appeal.

✅ Choose Qwen3.8 for

Value and potential. If you're budget-conscious and can wait for benchmarks, Qwen3.8's pricing advantage is enormous. Test it on your own workload through the Token Plan or Qoder. If it matches Qwen3.7's quality at similar pricing, the value proposition is unmatched.

What to Watch Next

Five specific things need to land before Qwen3.8 can be fairly compared to Kimi K3:

  1. Official benchmark table — Same detailed launch post with full benchmark tables that Qwen3.7 and Qwen3.6 received
  2. Active parameter count — The single most important number for understanding serving cost and self-hosting feasibility
  3. HuggingFace repo with license — If the open-weight promise is real, the repo must exist with a verifiable license file
  4. Published API pricing — To confirm Alibaba holds its value position rather than following Moonshot's premium pivot
  5. Independent third-party benchmarks — Artificial Analysis or similar, because launch numbers from any lab are marketing until reproduced

📅 Key date: July 27, 2026

Moonshot's open-weight deadline for Kimi K3. This is the first verifiable milestone in the Qwen3.8 vs K3 story. Watch for it closely.