⚡ TL;DR
- 🚀 Kimi K3 (Moonshot AI, July 16) has published benchmarks across 12 categories, 2.8T parameters, 16/896 MoE, and $3/$15 API pricing
- 📊 Qwen3.8-Max-Preview (Alibaba, July 19) claims 2.4T parameters and "second only to Fable 5" — but has zero published benchmarks
- 💰 Qwen's pricing edge is its biggest known advantage: $1.25 input / $3.75 output (Qwen3.7 baseline) vs K3's $3 / $15
- 🔓 Both promise open weights. K3's deadline is July 27. Qwen3.8's is "soon" — unconfirmed
- 📈 Kimi K3 leads Qwen3.7-Max on BenchAlign by 8.1 points (80.96 vs 72.84), with a 19.8-point gap in agentic tasks
- ⚠️ Direct Qwen3.8 vs K3 comparison is currently impossible — Qwen3.8 benchmarks don't exist yet
The Contenders
Kimi K3 — Moonshot AI
Released July 16, 2026, Kimi K3 is Moonshot AI's flagship open frontier model. At 2.8 trillion total parameters with only 16 of 896 experts active per token, it uses extreme sparsity (less than 2% active) to deliver frontier performance. K3 is available now via the Kimi app, Kimi Code CLI, API, and OpenRouter. Moonshot promises open weights by July 27, 2026. Verified: Moonshot official blog
Qwen3.8-Max-Preview — Alibaba
Announced July 19, 2026, Qwen3.8-Max-Preview is Alibaba's latest Qwen flagship. The Qwen team claims 2.4 trillion parameters and says the model is "second only to Fable 5" among frontier models. It's available through Alibaba's Token Plan ($6/$18/$68/month), Qoder, and QoderWork. Alibaba promises open weights "soon" but has published nothing yet. Unverified: Alibaba claims only
🎯 Why this matters right now
Both models launched within 3 days of each other in July 2026. Both are Chinese labs pushing past the trillion-parameter threshold. Both promise open weights. Both are positioned as alternatives to Claude Fable 5 and GPT 5.6 Sol. The timing is no coincidence — this is a direct competitive confrontation.
Architecture: Published vs Speculated
Kimi K3 Architecture
Moonshot published an extensive architecture blog for K3, detailing several novel components:
- Kimi Delta Attention (KDA) — Efficient attention designed for scaling sequence length and model depth
- Attention Residuals (AttnRes) — Selective retrieval mechanism across model depth
- Stable LatentMoE — Routing optimization at extreme sparsity (16/896 = 1.8% active), using quantile balancing with no heuristic updates and per-head Muon optimizer
- Sigmoid Tanh Unit (SiTU) — Activation control for training stability
- Gated MLA — Attention selectivity improvements
- Quantization-aware training — MXFP4 weights, MXFP8 activations for efficient serving
- 2.5x scaling efficiency improvement over Kimi K2
Qwen3.8 Architecture
Alibaba has published no architecture details for Qwen3.8. Based on the established Qwen lineage, we can infer:
- MoE — Virtually certain; every Qwen Max flagship since Qwen3-Max uses Mixture of Experts
- Hybrid attention — Likely; Qwen3-Next-80B-A3B combined transformer attention with linear recurrence
- GQA — Confirmed pattern across Qwen models (96 query heads, 8 key-value heads)
- Active parameters — Unknown; critical missing number that determines serving cost
- Expert count — Likely 160+ experts with 8-16 active per token, similar to K3
⚠️ Transparency gap
Kimi K3's architecture is fully documented in Moonshot's blog. Qwen3.8's architecture is entirely speculative. This asymmetry extends to every aspect of the comparison — K3 publishes, Qwen3.8 claims.
Benchmarks: Published vs Claimed
This is where the comparison gets interesting. Kimi K3 has published benchmark tables across 12 categories. Qwen3.8 has published zero benchmarks. What follows is the best available comparison: K3's published data vs Qwen3.7-Max (the previous generation, verified), with Qwen3.8's position estimated based on lineage.
Reasoning & Knowledge
Both models excel here, but K3 has the edge on reasoning benchmarks. The gap narrows on pure knowledge tasks where Qwen3.7-Max remains competitive.
| Benchmark | K3 (max) | Qwen3.7-Max | Qwen3.8 Claim |
|---|---|---|---|
| GPQA-Diamond | 93.5 | 92.4 | — |
| HLE (w/ tools) | 56.0 | 53.5 | — |
| Artificial Analysis Intelligence Index | 57.1 | 46.0 | — |
Coding & Agentic Tasks
K3's agentic coding benchmarks are strong, but Qwen3.7-Max has SWE-bench Verified data that K3 hasn't published on.
| Benchmark | K3 (max) | Qwen3.7-Max | Qwen3.8 Claim |
|---|---|---|---|
| Terminal-Bench 2.1 | 88.3 | 69.7 | — |
| SWE-bench Verified | — | 80.4 | — |
| FrontierSWE | 81.2 | — | — |
| SWE Marathon | 42.0 | — | — |
Agentic & Knowledge Work
K3's biggest advantage is in agentic tasks — a 19.8-point gap over Qwen3.7-Max on BenchAlign's Agentic category.
| Benchmark | K3 (max) | Qwen3.7-Max | Qwen3.8 Claim |
|---|---|---|---|
| BrowseComp | 91.2 | — | — |
| Toolathlon-Verified | 73.2 | — | — |
| AA Briefcase (Elo) | 1548 | 908 | — |
| GDPval-AA v2 (Elo) | 1668 | 1273 | — |
Multimodal & Vision
Qwen3.8 claims to be the first Qwen model over 1T parameters with multimodal support (images, videos, documents). K3 has native vision with strong MMMU-Pro and CharXiv scores.
| Benchmark | K3 (max) | Qwen3.7-Max | Qwen3.8 Claim |
|---|---|---|---|
| MMMU-Pro | 81.6 | — | Claimed |
| CharXiv | 84.8 | — | — |
| MathVision | 94.3 | — | — |
📊 BenchLM comparison (K3 vs Qwen3.7-Max)
BenchAlign aggregate: K3 80.96 vs Q3.7-Max 72.84 (8.1 point margin). K3 leads on Agentic (89.5 vs 69.7) and Knowledge (61.0 vs 64.2 — Qwen3.7 edge). Qwen3.7 leads on Coding, Reasoning, Math, and Multilingual — categories where K3 has no published data yet. Qwen3.8's position vs K3 cannot be measured until benchmarks are published.
Key Benchmark Gaps (K3 vs Qwen3.7-Max)
Pricing
Pricing is Qwen's strongest known advantage. Even the previous generation (Qwen3.7-Max) is dramatically cheaper than K3. Qwen3.8's pricing is not yet published, but Alibaba has consistently held its value position across generations.
💰 Pricing Comparison
| Model | Input / MTok | Output / MTok | Cache Hit | Pricing Tier |
|---|---|---|---|---|
| Qwen3.7-Max | $1.25 | $3.75 | — | Value |
| Kimi K3 | $3.00 | $15.00 | $0.30 | Premium |
| Claude Fable 5 (Sonnet) | $10.00 | $50.00 | — | Flagship |
✅ Qwen's pricing story
Qwen3.7-Max delivers a 92.4 GPQA score at $1.25 input — one eighth of Claude Fable 5's input price and less than half of Kimi K3's. If Qwen3.8 holds this pricing, it will be one of the best value models on the market regardless of where it ranks on benchmarks.
Open Weight Status
| Model | Status | Details |
|---|---|---|
| Kimi K3 | Promised: July 27, 2026 | Deadline set. Community watching. Moonshot has a track record of hitting dates. |
| Qwen3.8-Max-Preview | Promised: "Soon" | No date, no license, no HF repo. Would break pattern — last 2 Max flagships were closed. |
⚠️ Open-weight credibility check
Both Kimi K3 and Qwen3.8 promise open weights, but K3 has a specific date (July 27) while Qwen3.8 says only "soon." Critically, Alibaba's last two Max-tier flagships (Qwen3.7-Max and Qwen3.6-Max-Preview) both shipped closed through Alibaba Cloud Model Studio. An open-weight Qwen3.8 would break a clear pattern. Treat any promised open-weight release as unconfirmed until the HuggingFace repo exists with a real license file.
Scorecard
Confidence Ratings
Verdict
Current State: Kimi K3 Wins on Evidence, Qwen3.8 on Potential
As of July 20, 2026, Kimi K3 is the more complete product: published benchmarks across 12 categories, documented architecture, accessible API, and a specific open-weight deadline. Qwen3.8-Max-Preview is a promise wrapped in a teaser — the product exists (you can buy access), but every performance claim rests on Alibaba's word alone.
The Qwen3.8 team's claim of "second only to Fable 5" is not verifiable today. It may be true on some benchmarks — Qwen3.7-Max already scores 92.4 on GPQA-Diamond, within 1.1 points of K3's 93.5 — but closing a 15-point gap on SWE-bench Verified (Fable 5 at 95% vs Qwen3.7 at 80.4%) in one generation would be extraordinary.
Qwen3.8's real weapon is pricing. If it delivers even marginally better benchmarks than Qwen3.7-Max while holding the $1.25/$3.75 pricing tier, it will be the best value model on the market regardless of where it ranks on leaderboards.
Who Should Use Which?
Maximum capability today. If you need proven performance on coding, reasoning, agentic tasks, and multimodal work — and you're willing to pay premium pricing — K3 is the stronger model right now. Its open-weight release on July 27 will further expand its appeal.
Value and potential. If you're budget-conscious and can wait for benchmarks, Qwen3.8's pricing advantage is enormous. Test it on your own workload through the Token Plan or Qoder. If it matches Qwen3.7's quality at similar pricing, the value proposition is unmatched.
What to Watch Next
Five specific things need to land before Qwen3.8 can be fairly compared to Kimi K3:
- Official benchmark table — Same detailed launch post with full benchmark tables that Qwen3.7 and Qwen3.6 received
- Active parameter count — The single most important number for understanding serving cost and self-hosting feasibility
- HuggingFace repo with license — If the open-weight promise is real, the repo must exist with a verifiable license file
- Published API pricing — To confirm Alibaba holds its value position rather than following Moonshot's premium pivot
- Independent third-party benchmarks — Artificial Analysis or similar, because launch numbers from any lab are marketing until reproduced
📅 Key date: July 27, 2026
Moonshot's open-weight deadline for Kimi K3. This is the first verifiable milestone in the Qwen3.8 vs K3 story. Watch for it closely.