2026 DEEPSEEK V4
FLASH_0731_
OFFICIAL.
TL;DR upfront: 31 июля DeepSeek выпустил V4-Flash-0731 — не новую архитектуру, а тот же 284B MoE с перезапуском post-training. Agent-бенчмарки обгоняют V4-Pro preview (1,6T/49B active), API стоит долю от Claude Opus 4.8. V4-Pro GA и Harness — «скоро», без даты. Ниже: полный timeline апрель–август, pricing matrix, CSA+HCA/mHC/Muon stack, Harness caveats, cross-bench vs Kimi K3/GLM-5.2/Qwen3.8-Max, controversy flags, «kill line» industry read, 5-step integration runbook, FAQ.
30 секунд · TL;DR
| Релиз | 2026-07-31 V4-Flash-0731 официальная API beta · только API, app/web не обновлены |
| Архитектура | 284B / 13B active, идентична april preview · прирост только post-training |
| Цена | Input $0,14 (cache miss) / $0,0028 (hit) · Output $0,28/M |
| Agent | Terminal Bench 2.0 82,7 (vendor + Harness minimal mode) · выше V4-Pro preview 67,9 |
| Ожидается | V4-Pro official · agent framework Harness · дата не подтверждена |
1. Почему 0731 — не cosmetic update
- Pro flagship задержан, Flash обогнал preview. Community называла Liang Wenfeng «Liang Baikai» — Pro GA обещали на середину июля. 31 июля Flash 0731 с меньшим param budget бьёт Pro preview на agent scores. Pro + Harness — no ETA.
- Benchmarks = harness-dependent. Terminal Bench 2.0 (82,7) измерен через неопубликованный Harness minimal mode. Vendor disclaimer: «scores extremely sensitive to harness choice» — не переносите 1:1 в Claude Code/Cursor.
- API-only rollout. chat.deepseek.com не обновлён — devs и end users на разных build tracks. Expected behavior, не баг.
2. Timeline: april preview → july «official»
| Дата | Событие |
|---|---|
| 2026-04-24 | DeepSeek-V4 preview + open weights (MIT): V4-Pro (1,6T/49B active), V4-Flash (284B/13B active), 1M context window |
| 2026-07-24 | Legacy aliases deepseek-chat / deepseek-reasoner EOL; traffic → V4 naming |
| 2026-07-27 | Moonshot AI Kimi K3 (2,8T) full weights на Hugging Face — крупнейший open-weight drop года |
| 2026-07-31 | V4-Flash-0731 official API beta; weights HF; changelog первый раз называет «DeepSeek Harness» |
| На 2026-08-05 | V4-Pro official — «as soon as possible». Китайские медиа: GA window 10–20 августа — не подтверждено DeepSeek |
3. Pricing & spec matrix
| Модель | Статус | Total / Active | Context | Input (miss/hit, $/M) | Output ($/M) | License |
|---|---|---|---|---|---|---|
| V4-Flash-0731 | Official (31.07) | 284B / 13B | 1M | $0,14 / $0,0028 | $0,28 | MIT |
| V4-Pro | Preview (24.04, official pending) | 1,6T / 49B | 1M | $0,435 / $0,003625 | $0,87 | MIT |
| Kimi K3 | Open weights (27.07) | 2,8T / ~104B (community est.) | ~1,05M | $3,00 / $0,30 | $15,00 | Kimi K3 License |
| GLM-5.2 | Open (июнь 2026) | ~744B / ~40B | 1M | Не верифицировано | Не верифицировано | MIT |
| Qwen3.8-Max | API GA (02.08), weights pending | 2,4T / 95B | 1M | $2,00 / ~$0,17–0,25 | $6,00 | Open-source promised |
Все цены — vendor list prices. DeepSeek анонсировал 2× peak surcharge (Пекин 9–12, 14–18) — effective date TBD.
4. Technical deep-dive: откуда performance
4.1 Architecture frozen — post-training rerun
V4-Flash-0731 byte-for-byte same param topology как april preview. Vendor statement: весь agent score jump = fresh post-training pass, zero scale-up. 284B/13B beating 1,6T/49B из той же family — signal: H2 2026 competition shifting от pure param count к post-training data quality + RLHF/RLAIF pipeline craft.
4.2 Hybrid attention stack: CSA+HCA, mHC, Muon
Per technical report «DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence»:
- Hybrid attention: Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA), branded «DSA Sparse Attention» — cuts compute + VRAM at 1M context;
- Manifold-Constrained Hyper-Connections (mHC): upgraded residual path;
- Muon optimizer: faster convergence, training stability at 61 layers / 1,6T scale.
Vendor claim @ 1M tokens: V4-Pro = 27% per-token FLOPs vs V3.2, KV cache footprint 10%. Independent reproduction pending — если подтвердится, million-token context становится economically servable, не marketing checkbox.
4.3 Harness: first-party agent execution framework
31 июля — first public mention DeepSeek Harness: file I/O, tool calls, multi-step engineering tasks. Positioning vs Claude Code (которым DeepSeek teams пользовались до сих пор). Все published agent benchmarks (Terminal Bench 2.0, Toolathlon) = Harness minimal mode (max effort, top_p=0.95, temp=1.0), framework not yet released. Treat scores as «vendor + specific harness», not portable capability claim.
5. Cross-bench: китайский open-weight cluster
| Модель | Lab | Release | Total | AA Intelligence Index | Cost/task (AA) |
|---|---|---|---|---|---|
| V4-Flash-0731 | DeepSeek | 31.07 official | 284B | 50 | $0,03 |
| Kimi K3 | Moonshot AI | 16.07 preview / 27.07 weights | 2,8T | 57 | $0,86 |
| GLM-5.2 | Zhipu/Z.ai | Июнь 2026 | ~744B | ~1 pt above Flash | Не верифицировано |
| Qwen3.8-Max | Alibaba | 02.08 GA | 2,4T | Не верифицировано | Не верифицировано |
| GPT-5.6 Sol | OpenAI | Closed | — | 9+ pts above | $1,86 |
| Claude Fable 5 | Anthropic | Closed | — | 9+ pts above | $3,15 |
Intelligence Index + cost/task: Artificial Analysis (independent). DeepSeek agent scores: vendor-reported — methodology mismatch, не суммируйте.
Key tension: V4-Flash не лидер AA index (trails Kimi K3, GLM-5.2), но cost/task ≈ 1/29 K3, 1/62 GPT-5.6 Sol, 1/105 Claude Fable 5. DeepSeek optimizes «good-enough intelligence @ floor price» — explains 7-week OpenRouter #1 streak на preview build.
6. Caveats: benchmarks ≠ ground truth
- Harness lock-in, vendor-confirmed. Terminal Bench 2.0 (82,7 vs 67,9) = unreleased framework + self-reported.
- Production friction reports. 21st Century Business Herald: low cache-hit rates, safety classifier timeouts — compute budget still caps what post-training alone fixes.
- Pro/Harness dates unconfirmed. «Aug 10–20 GA» = unnamed media sources; official changelog: «as soon as possible» only.
- Funding/IPO rumors. ~$7,4B round, ~$48,7B valuation — «according to reports», zero regulatory filing.
7. 5-step integration runbook
- Migrate model IDs: grep
deepseek-chat/deepseek-reasoner→deepseek-v4-flashordeepseek-v4-pro(EOL since 24.07). - Flash vs Pro by workload: batch agents, routing, high-QPS → Flash 0731; deep reasoning → Pro preview, re-eval on official GA.
- Prompt cache для cost floor: repeat system prompts; hit cost $0,0028/M — near-zero marginal.
- Peak surcharge scheduling: non-real-time batch jobs до 9:00 или после 18:00 Beijing.
- Mac-side validation: Cursor/OpenClaw Flash 0731 vs Kimi K3/Qwen3.8-Max на identical SWE-bench subset — measure real latency + pass rate, ignore vendor leaderboards.
8. FAQ
Q: V4 vs V3.2 — главное отличие?
A: Native 1M context; FLOPs 27%, KV cache 10% vs V3.2 (vendor); agent tuning для Claude Code, OpenCode, etc.
Q: Flash или Pro daily?
A: Batch, agent pipelines, cost-sensitive → Flash 0731; deep reasoning + budget → Pro preview.
Q: V4-Pro official доступен?
A: Нет на 05.08.2026. Только Flash 0731, API-only. «Aug 10–20» = rumor.
Q: Верить benchmark numbers?
A: SWE-bench Verified (third-party) — relatively solid. Terminal Bench 2.0 Harness-locked — ждите community repro с Claude Code/Cursor.
Q: Практический impact?
A: OpenAI/Anthropic-compatible API; deepseek-v4-flash → new official build; migrate legacy aliases ASAP. MIT = self-host path open.
9. Industry read: «kill line» и price war escalation
Pre-0731: «Liang Baikai» mockery за Pro delay. Post-release: «Liang Sheng» — sentiment barometer китайской AI dev community.
Substantive frame: 「斩杀线」 (kill line) — DeepSeek sets floor: adequate perf + rock-bottom price. Competitors без clear capability edge и без lower price → market irrelevance. Explains GPT-5.6 Luna ~80% price cut same window. Incubator quote (21st Century Business Herald): «Every LLM company runs ahead of Liang Wenfeng — must stay ahead to survive.»
31 июля: Nvidia, Broadcom, AMD flat — market normalized «DeepSeek efficiency breakthrough», unlike R1 chip selloff Q1 2025.
Mac dev angle: MIT + 13B active → quantized V4-Flash experiments на Apple Silicon unified memory feasible. Community MLX quant runs ongoing; track Apple Silicon V4 AMX-2 benchmarks для local tokens/sec и Metal throughput baselines. Production agent pipelines + SWE-bench regression — isolated environment; не забивайте unified memory, иначе Cursor/Xcode thrash на swap.
10. Closing: API anywhere — agent validation on Mac/Metal
OpenRouter route swap, API key paste — Windows/Linux достаточно для API routing. Но Cursor + V4-Flash long-context coding, OpenClaw multi-channel agent A/B, MLX quantized Flash validation перед prod deploy — Apple Silicon unified memory + Metal compute stack = lowest-friction path. Zero PCIe bottleneck на KV-heavy workloads vs discrete GPU + system RAM split; Metal Performance Shaders path даёт predictable throughput на batch inference.
Pragmatic split: daily driver на API, Flash 0731 batch agent stress, Harness regression post-release, 1M context doc batch jobs — на MACGPU remote Mac mini M4 node. On-demand, SSH-secured, peak-valley cron на remote silicon. Main machine stable для everyday dev; remote Metal node съедает agent memory pressure и даёт reproducible throughput numbers.