OPUS 5
ПОЛЦЕНЫ_
FLAGSHIP · K3 СКАНДАЛ.
Резюме: Две новости этой недели выглядят несвязанными, но сходятся в одном — кто даёт frontier-интеллект по цене, которую разработчик может позволить, и откуда эта мощность реально берётся. Боль: выбирать между daily-driver Opus 5 от Anthropic и ultra-cheap open-weight K3 от Moonshot, пока обвинения Белого дома в дистилляции и K3, называющий себя Claude, мутят воду. Вывод: Opus 5 — прагматичный апгрейд для compliance-sensitive workflow; спор вокруг K3 не закрыт, но deployment-ID leak Greenblatt — самый воспроизводимый технический сигнал. Структура: разбор Opus 5 → таблица Opus vs Fable → таймлайн дистилляции K3 → «почему K3 говорит Claude» → 5 шагов Mac routing → индustry analysis → FAQ.
TL;DR за 30 секунд
| Opus 5 | Релиз 24.07 · дефолт Claude Max · $5/$25 за M токенов · gap 0,5% vs Fable 5 на CursorBench |
| K3 | Launch 16.07 · 2,8T params · веса 27.07 · обвинение Белого дома 22.07 |
| Технический hook | K3 self-identify как Claude · эмитит claude-opus-4-5-20250929 |
| Контраргумент по timeline | Fable 5 публичен с 1.07 · 2 недели до K3 · эксперты скептичны |
1. Pain points: почему обе истории критичны сейчас
- Цена flagship vs gap daily-driver. Fable 5 лидирует в бенчмарках, но ~$10/$50 за миллион токенов и 30-day data retention opt-in. Pro/Max ждали «half-price flagship» — Opus 5 ответ Anthropic.
- Open weights догоняют, provenance под вопросом. K3 на графиках Moonshot уступает только Fable 5 и GPT-5.6 Sol — за долю цены. Когда Белый дом говорит об «industrial-scale distillation», compliance и supply-chain risk становятся реальными критериями выбора.
- Политическая риторика vs воспроизводимые сигналы. X-post Kratsios без публичных доказательств. Community встреpеножило: Greenblatt показал, что K3 выдаёт Anthropic deployment metadata точнее, чем сами Claude-модели о себе.
- Давление на Mac routing stack. Multi-model стеки Cursor и OpenClaw должны совместить «compliant Claude API» с «дешёвым open-weight fallback». См. наш гайд OpenRouter и countdown весов K3 — события недели бьют прямо в routing tables.
2. Claude Opus 5: самый практичный модель Anthropic
24 июля 2026 Anthropic выкатил Claude Opus 5 (model ID: claude-opus-5) и сразу сделал его дефолтом Claude Max плюс сильнейшей моделью для Claude Pro. Pricing без изменений vs Opus 4.8: $5 / $25 за миллион input/output токенов. Context window: 1M tokens (единственный tier). Max output: 128K tokens. Thinking включён по умолчанию.
2.1 Скачок бенчмарков при flat pricing
| Benchmark | Opus 5 | Комментарий |
|---|---|---|
| Frontier-Bench v0.1 | Обходит все модели | 2×+ vs Opus 4.8 при меньшем cost-per-task |
| CursorBench 3.2 (max effort) | 0,5% ниже пика Fable 5 | Половина cost-per-task |
| ARC-AGI 3 | 3× следующей модели | Скачок novel problem-solving |
| OSWorld 2.0 | Обходит best Fable 5 | Менее 1/3 cost Fable 5 |
| Zapier AutomationBench | 100% pass rate | Ни одна prior model не закрыла workflow |
Команда Cursor: «Claude Opus 5 даёт near-Fable-5 intelligence при Opus speed и cost.» Box: +11% data-analysis workflows, +17% due diligence, +8% overall accuracy.
2.2 Safety и data retention
Opus 5 — наиболее aligned модель Anthropic: минимальный deceptive-behavior rate, сложнее trick в misuse. Но Anthropic сознательно не выводил Opus 5 на frontier dual-use cyber/bio risk — это lane Mythos 5. Cyber classifiers вмешиваются ~85% реже, чем у Fable 5: source-code vulnerability discovery возможен, binary scanning / pentesting / exploit generation заблокированы.
Легко пропустить: для general Opus 5 access нет forced data retention. Fable 5 и Mythos 5 требуют 30-day retention opt-in — значимый compliance differentiator.
3. Claude Opus 5 vs Fable 5: decision matrix
| Параметр | Claude Opus 5 | Claude Fable 5 |
|---|---|---|
| Input / output pricing | $5 / $25 за M tokens | ~$10 / $50 |
| CursorBench 3.2 peak | 0,5% ниже Fable 5 | Текущий coding ceiling |
| Claude Max default | Да (с 24.07) | Нет |
| Data retention | Не требуется по умолчанию | 30-day opt-in |
| Dual-use risk frontier | Сознательно capped | Сильнее (Mythos 5 restricted) |
| Best for | Daily coding, agents, compliance workflows | Peak benchmark, absolute frontier |
4. Спор дистилляции Kimi K3: Белый дом vs timeline math
Moonshot AI выпустил Kimi K3 16 июля 2026: 2,8 триллиона total params, sparse MoE (16 из 896 experts active, ~50B active-parameter equivalent), 1M-token context с native vision. GPQA-Diamond 93,5%, BrowseComp 91,2%. Full weights обещаны на 27 июля — externally unverifiable во время скандала.
4.1 Обвинение Белого дома и prior claims Anthropic
22–23 июля OSTP Director Michael Kratsios заявил о «large-scale, covert industrial distillation» Fable и отдельно — export-restricted Nvidia GB300 через Thailand. Treasury Secretary Scott Bessent упомянул «watermarks» US LLM на китайских моделях без определения термина.
Контекст: в феврале 2026 Anthropic публично назвал Moonshot, DeepSeek и MiniMax, заявив 3,4M+ anomalous API interactions как «deliberate capability extraction».
4.2 Independent researchers спорят timing
«Fable публичен только с 1 июля. Нельзя destill столько данных, обучить модель и выпустить за две недели.» — Braden Hancock, co-founder Snorkel AI
Nathan Lambert (Allen Institute for AI): по мере приближения китайских моделей к frontier и сдвига training к RL, простая distillation даёт diminishing returns — будь это easy, все бы уже догнали.
5. Почему Kimi K3 упорно называет себя Claude
Это angle, который стоит времени technical reader — и почти никто вне GitHub Greenblatt (rgreenblatt/which_claude_is_k3) не разобрал его глубоко.
Chief Scientist Redwood Research Ryan Greenblatt опубликовал statistical analysis ~24 июля: на identity probe Kimi K3 диспропорционально self-identify как Claude — и не vaguely. Иногда эмитит exact internal Anthropic deployment IDs:
Настоящие Claude-модели так о себе не говорят. Claude Sonnet 4.5: «I'm Claude Sonnet 4.5.» Opus 4.5 часто не даёт version string или ошибается.
Интерпретация Greenblatt: модель, воспроизводящая deployment metadata teacher точнее, чем teacher о себе, плохо объясняется conversational mimicry. Указывает на training на Claude data с deployment metadata labels — API logs или synthetic data с metadata tags — specific, harder-to-wave-away form distillation.
Деталь: leaked identity K3 указывает на «Claude 4.5 era» (late 2025), не текущую Fable/Mythos generation — Kimi K2 signal был на earlier Claude Sonnet 4. Pattern «generation-by-generation catch-up».
Caveat: Greenblatt подчёркивает — это не доказывает distillation; contamination, leaked system prompts или public synthetic datasets тоже возможны. Вместе с February accusation Anthropic — первый reproducible technical signal в saga, не political rhetoric.
6. Пять шагов: настроить Mac developer model stack
- Compliance first. Enterprise data или export-control-sensitive work → no-default-retention Opus 5 + official API beats K3, пока provenance не cleared.
- Update Cursor / OpenClaw primary model. Claude Max users default на Opus 5. Explicit benchmark
claude-opus-5vs Fable 5 на real project cost-per-task. - Configure OpenRouter fallback chain. Primary Opus 5 → rate-limit fallback DeepSeek V4 Flash или K3 API (если provenance risk acceptable) → free tier для prompt debugging по гайду OpenRouter.
- Calendar 27.07 для K3 weights. После release — isolated benchmarks и identity-probe tests на non-production node; reproduce Greenblatt statistics самостоятельно.
- Isolate stress tests от daily dev. Long-context agent runs и multi-model A/B не должны saturate unified memory MacBook — offload на remote Mac node (см. closing).
7. Timeline
| Дата | Событие |
|---|---|
| Фев 2026 | Anthropic обвиняет Moonshot/DeepSeek/MiniMax в industrial distillation |
| 1 июл | Claude Fable 5 publicly available |
| 16 июл | Kimi K3 API/product launch |
| 22–23 июл | Белый дом: Kratsios distillation + chip allegations |
| 24 июл | Claude Opus 5 release; Greenblatt identity analysis |
| 27 июл (planned) | Kimi K3 full open weights |
8. Industry analysis: price war meets provenance
Opus 5 и K3 row вместе — phase shift: frontier capability commoditizing. Anthropic не discounting old inventory — restructuring product line: Fable 5 держит absolute frontier, Opus 5 daily driver, Mythos 5 locks dual-use risk.
Moonshot бьёт в тот же market open weights, ultra-low price, fewer refusals. K3 controversy — первый public flashpoint более ugly industry question: когда lab заявляет frontier performance за 1/10 cost — как отличить better engineering от quietly riding чужой model?
Для Mac developers это не spectator sport. Cursor defaults, OpenClaw fallback chains, OpenRouter routing tables перепишутся за 30 дней: Opus 5 default, K3 weights drop. Deployment-ID angle Greenblatt важен, потому что даёт reproducible, quantifiable detection — run identity probes yourself; не ждите government investigations.
r/LocalLLaMA splits три ways: excitement — open-closed gaps «days not months»; jokes — nobody runs 2.8T locally; pragmatists — real sell K3 это price и lack of refusals, не «beating Fable 5». До 27.07 architecture claims остаются «vendor self-report + external guesswork».
На уровне inference stack: unified memory Apple Silicon + Metal toolchain даёт predictable throughput для long-session agents без CUDA memory fragmentation. Для identity-probe reproduction и weight benchmarks K3 изолированный remote node с SSH — единственный sane path, если daily driver не 128GB+ конфиг.
9. FAQ
Q: Насколько Opus 5 дешевле Fable 5?
A: ~half per token ($5/$25 vs ~$10/$50), within 0,5% CursorBench peak.
Q: Opus 5 — default Claude Max?
A: Да, с 24.07.2026.
Q: Moonshot distill K3 из Claude?
A: Unconfirmed. Белый дом без public evidence; timeline contested; Greenblatt Claude self-ID — strongest technical indirect evidence.
Q: Когда full weights K3?
A: Committed 27.07.2026 — not yet at publication.
Q: Почему K3 говорит Claude?
A: Statistical bias + internal deployment ID strings — likely training-data contamination с metadata-tagged Claude samples; not conclusive distillation proof.
10. Closing: Opus 5 для compliant API, Mac nodes для stress tests
Claude API или Kimi K3 на Windows/Linux cloud boxes — ок для daily coding. Но Cursor + Opus 5 long-session agents, OpenClaw multi-model fallback stress tests или isolated K3 identity-probe reproduction на Mac требуют Apple Silicon unified memory + Metal toolchain + stable 24/7 — минимальный friction path для throughput-bound workloads.
Practical split: primary machine на Opus 5 / compliant OpenRouter routing; offload agent stress tests, K3 weight benchmarks, long-context batches на MACGPU remote Mac mini M4 nodes — rent on demand, SSH-isolated, чтобы 2.8T-parameter experiments не brick daily driver. Half-price flagship и disputed open frontier coexisting this week — isolated validation beats betting на one stack.