GROK 4.6
AUG_7_
1.5T_
TARGET.
Elon Musk says Grok 4.6 is coming "around August 7," with a larger Grok 4.7 to follow a few weeks later — shared in an X reply to Vercel CEO Guillermo Rauch on July 28, 2026, not in an official xAI announcement. Pain point: developers face "Musk time" schedule elasticity, zero benchmark data on announced specs, and a selection window compressed by Kimi K3's open-weight shock. Conclusion: until xAI ships a model card, token efficiency and real task cost — not parameter count alone — are the durable decision basis. Structure: timeline, spec tables, SFT/RL breakdown, competitive comparison, caveats, August release-bottleneck analysis, five-step guide, FAQ.
30-second executive summary
| Source | Musk's July 28 X reply (not an xAI blog post) |
| Grok 4.6 | ~Aug 7 · 1.5T params · SFT/RL upgrade focus |
| Grok 4.7 | Late Aug–early Sep · 2.1T · better than 4.6, slightly slower to serve |
| Competitive context | ~10 days after Kimi K3 open weights · rumored Fable 5.1 same month |
| Pricing / benchmarks | All blank — no model card or third-party scores |
1. Pain points: why this preview matters — and why you cannot trust it fully
- "Musk time" has a track record. xAI, Tesla, and SpaceX timelines from Musk have historically slipped by days to weeks.
- The only source is a tweet. Unlike Grok 4.5's launch with 15 tracked benchmarks and a published model card, 4.6/4.7 have zero independent verification.
- Kimi K3 raised the competitive bar. K3's July 26 open-weight release (2.8T, Frontend Code Arena 1,679 points) drew praise from Musk himself; compressing xAI's release cadence is the most plausible explanation for the accelerated roadmap.
2. Timeline: Grok 4.5 → 4.6 → 4.7
| Date | Event |
|---|---|
| Jul 8, 2026 | xAI ships Grok 4.5: coding/agentic flagship, co-trained with Cursor, 500K context, $2/$6 per 1M input/output tokens |
| Jul 16 | Moonshot AI's Kimi K3 hosted preview goes live |
| Jul 26 | Kimi K3 full open-weight release (2.8T, 1M context) |
| Jul 28 | Musk posts Grok 4.6/4.7 roadmap; same day 1,200+ employees publish "Pacing the Frontier" letter |
| ~Aug 7 | Grok 4.6 target: 1.5T params, SFT/RL upgrade focus |
| ~Late Aug–early Sep | Grok 4.7 target: 2.1T params, broadly better than 4.6, slightly slower serving, better token efficiency |
3. Core data table
| Model | Date | Parameters | Focus | Status |
|---|---|---|---|---|
| Grok 4.3 Beta | Apr 17, 2026 | Undisclosed | Baseline | Shipped |
| Grok 4.5 | Jul 8, 2026 | Undisclosed | Coding/agentic, Cursor co-training | Shipped, benchmarked |
| Grok 4.6 | ~Aug 7 | 1.5T | SFT/RL upgrade | Announced via tweet |
| Grok 4.7 | ~late Aug–early Sep | 2.1T | Broad upgrade over 4.6 | Announced via tweet |
4. Why xAI is emphasizing post-training, not just scale
4.1 SFT and RL in plain terms
Supervised fine-tuning (SFT) shapes behavior with curated examples; reinforcement learning (RL) teaches multi-step agentic sequences via reward signals. Musk's "significantly improved SFT & RL" framing continues Grok 4.5's playbook: real Cursor developer sessions helped deliver Terminal-Bench 2.1 83.3%, SWE-Bench Pro 64.7%, with roughly 15,954 output tokens per SWE-Bench Pro task vs Opus 4.8's 67,020 — a 4.2× efficiency gap.
4.2 The scale-vs-speed trade-off
Grok 4.6 at 1.5T is a real scale jump, but Musk's Grok 4.7 framing — bigger at 2.1T, "better in every way except slightly slower to serve" — suggests two SKUs with different latency/quality trade-offs, consistent with Anthropic's Sonnet/Opus split and OpenAI's mini/full tiers.
4.3 Competitive pressure from Kimi K3
Grok 4.6's target lands almost exactly 10 days after K3's open-weight release rattled the industry. K3 topped Frontend Code Arena at 1,679 points — the first open-weight model to beat every closed model on that board — and ranked third on Artificial Analysis's Intelligence Index.
5. How Grok stacks up against the field
| Model | Vendor | Parameters | Context | Pricing (in/out per 1M) | Source |
|---|---|---|---|---|---|
| Grok 4.5 | xAI | Undisclosed | 500K | $2 / $6 | xAI official |
| Grok 4.6 (announced) | xAI | 1.5T | Undisclosed | Undisclosed | Musk X post (unverified) |
| Kimi K3 | Moonshot AI | 2.8T (MoE, ~16/896 active) | 1M | $0.30 cache hit / $3 miss in, $15 out | Moonshot official |
| Claude Fable 5.1 (rumored) | Anthropic | Undisclosed | Undisclosed | Rumored $10 / $50 | 36kr, WinCentral leaks |
| GPT-5.6 Sol | OpenAI | Undisclosed | Undisclosed | Undisclosed | OpenAI official |
6. What to flag before you trust this timeline
- Only source is a tweet — no xAI blog post, model card, or product page.
- Benchmarks and pricing are blank — "1.5T parameters" and "SFT/RL upgrade" are unverified vendor claims.
- xAI safety controversies — July 2026 lawsuit over CSAM generation via bypassed safeguards; January 2026 Common Sense Media rated Grok among the worst for child-safety risks.
- Industry split on AI pacing — same day Musk announced 4.6/4.7, 1,200+ employees at OpenAI, Anthropic, Google DeepMind, and Meta published "Pacing the Frontier"; xAI is absent from the signatory list.
7. Industry insight: August 2026 as a release bottleneck
If Musk's timeline holds, Grok 4.6 and 4.7 land in the same month as a rumored Claude Fable 5.1 and just weeks after Kimi K3's open-weight shock — making August 2026 one of the densest frontier-model release months on record. For teams evaluating models, the useful shelf life of any single flagship compresses to weeks, not quarters.
The deeper shift is from leaderboard rank to per-task total cost: Grok 4.5 proved output-token efficiency can offset lower intelligence-index scores; Kimi K3 made self-hosting a hard option; a rumored Fable 5.1 would further squeeze closed-model premiums. Three paths — closed API, open self-hosting, hybrid routing — face a concentrated stress test in August.
For engineering teams, the operational risk is routing-table drift: model aliases, fallback chains, and cost caps that were quarterly maintenance items become weekly incidents if three trillion-parameter-class models ship in the same 30-day window.
8. Five-step selection guide before the model card drops
- Lock your information sources: xAI blog, Grok Build console, Musk's X account — do not bake announced specs into SLAs pre-launch.
- Build a three-tier route: primary (current Grok 4.5 or Kimi K3 API) → high-precision review (Claude/GPT) → offline fallback (local MLX quantized open weights).
- Baseline token efficiency: record output tokens per real repo task now; run 20 production-like tasks on any new model before switching defaults.
- August release-window playbook: reserve a 48-hour evaluation window per major launch — no full production cutover on day one.
- Vendor risk audit: include xAI's recent safety litigation alongside technical benchmarks in enterprise reviews.
9. FAQ
When exactly is Grok 4.6 coming out?
Musk said "around August 7, 2026" in an X post; xAI has not officially confirmed a date.
What's the difference between Grok 4.6 and 4.7?
4.6 is 1.5T with SFT/RL focus; 4.7 is 2.1T, expected weeks later, outperforming 4.6 except on serving speed where it trades latency for better token efficiency.
Will Grok 4.6 beat Kimi K3 or Claude Fable 5.1?
Too early to tell — no published Grok 4.6 benchmarks; K3 has verified scores; Fable 5.1 is unconfirmed by Anthropic.
How much will Grok 4.6 cost?
Unknown. Grok 4.5 launched at $2/$6 per million input/output tokens as a reference point.
Where will I be able to use Grok 4.6?
Based on Grok 4.5's rollout: likely Grok Build, xAI API, and console first, then third-party platforms like Cursor.
10. Closing: frontier API chase + Mac local open-weight hedge
A Windows/Linux cloud box can track Grok 4.6 headlines and call APIs, but it falls short on local open-weight baselines, Cursor/OpenClaw agent residency, MLX-quantized Kimi K3 fallbacks, and unified-memory long-context routing. In an August release bottleneck, the scarce resource is not "which model scores highest" but auditable compute that can run 48-hour real-task comparisons without vendor lock-in. If you need local MLX for K3 quantization, 24/7 remote nodes for OpenClaw multi-model routing tests, and unified memory for agent session splitting, use a three-tier stack: local MLX for daily work and offline hedge; frontier closed APIs for Grok/Claude hard reasoning; MACGPU remote Mac nodes for agent residency and isolated evaluation — the best hedge against both "Musk time" and a release-month routing storm.