Our Blog

Blog Index

Grok 4.6 Arrives: SpaceXAI's New Frontier Model Targets Long-Running Agents and Coding

Posted on 13th Aug 2026 06:07:50 in Artificial Intelligence, Machine Learning

Tagged as: Grok, SpaceXAI, xAI, AI Models, AI Agents, Coding

SpaceXAI — the company formerly known as xAI — released Grok 4.6 on August 12, 2026, its latest frontier model and the direct successor to Grok 4.5, which launched barely a month earlier in July. The new model is built with a specific focus: long-running agents and ambitious interactive and visual work. Where earlier Grok releases were judged primarily on conversational ability and raw benchmark scores, Grok 4.6 is designed to stay with complex tasks across many steps — researching a topic, analyzing information, working across a codebase, or turning a broad product idea into a polished, working application. Independent evaluator Artificial Analysis places it as the world's third-best model on its composite Intelligence Index, matching OpenAI's GPT-5.6 Sol and overtaking Moonshot AI's open-weight Kimi K3, with only Anthropic's Claude Opus 5 and Fable 5 ahead.

Built for Work That Spans Days, Not Minutes

The headline change in Grok 4.6 is behavioral rather than architectural. SpaceXAI says the model underwent a longer supplemental training run than Grok 4.5, using curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. Grok 4.5 was then used to regenerate supervised fine-tuning trajectories across reasoning levels, agent harnesses, and domains including STEM, software engineering, and knowledge work, with problematic traces filtered out by model-based checks. Reinforcement learning targeted agentic environments spanning general coding, knowledge work, kernel optimization, web development, and computer-aided design.

The result, according to the company, is a model that verifies its own work: on longer trajectories, SpaceXAI observed more self-testing and verification, with the model checking its output before moving on. It also produces stronger first passes on visual and interactive projects than Grok 4.5, establishing structure and a visual language for an application in a single pass — useful for teams that want to start from a substantial scaffold and iterate in the loop. In testing, the model proved especially strong at turning a broad product idea into a working first version, researching unfamiliar domains as it goes.

Benchmarks: At the Frontier, But Not Sweeping It

Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks — matching GPT-5.6 Sol Max and improving five points over Grok 4.5 High, while Fable 5 Max leads at 62. Its strongest individual results come from longer-horizon professional work and agentic coding:

  • GDPVal-AA v2 (Elo 1,753): real-world tasks like scheduling and diagramming — ahead of GPT-5.6 Sol Max (1,728) and Fable 5 Max (1,741), and up sharply from Grok 4.5's 1,526.
  • AA-Briefcase (Elo 1,577): longer-horizon professional work, narrowly edging Fable 5 Max (1,574) and topping GPT-5.6 Sol Max (1,502).
  • CursorBench v3.2 (69.9%): coding-agent performance, up from 66.7% on Grok 4.5, behind only Fable 5 Max's 70.5%.
  • FrontierCode v1.1 Extended (61.3%): up from 56.6%, ahead of GPT-5.6 Sol Max's 60.6% but behind Fable 5 Max's 63.6%.
  • APEX-Agents (57.5%): a 10.4-point jump over Grok 4.5's 47.1%, narrowly exceeding GPT-5.6 Sol Max's 56.7%.
  • Harvey LAB (15.8%): legal-agent evals — well ahead of GPT-5.6 Sol Max (2.5%) and Fable 5 Max (11.3%).

There are gaps too. On DeepSWE v1.1, Grok 4.6 reaches 65.9% — a solid gain over Grok 4.5's 54% — but trails GPT-5.6 Sol Max (73%) and Fable 5 Max (70%). Terminal-Bench v3.0 exposes the largest remaining gap: 26% versus 34.6% for GPT-5.6 Sol Max and 34.1% for Fable 5 Max. In other words, Grok 4.6 has reached the frontier without establishing an uncontested lead, and its agentic-efficiency claims will need to survive production workloads.

Pricing and Availability: Mid-Priced Frontier

Grok 4.6 is available immediately in Cursor (desktop, cloud agents, iOS, CLI, and SDK), in Grok Build — SpaceXAI's answer to Claude Code and Codex, included in the $30-per-month SuperGrok plan — and through the xAI API, plus partners including OpenRouter, Vercel, and Cloudflare. API pricing starts at $2 per million input tokens and $6 per million output tokens, with cached input at $0.50 per million. Requests exceeding a 200K-token prompt pay the higher-context rate of $4/$12, and a fast variant runs at twice the standard price. SpaceXAI is offering double the included usage inside Grok Build and Cursor for the first week.

At $8 per million tokens for a combined in/out pass, Grok 4.6 is a mid-priced frontier model: less than half of GPT-5.6 Sol in standard mode ($5/$30), well under Claude Fable 5's $10/$50, and cheaper than Moonshot's Kimi K3 at $3/$15 — the open-weight Chinese model it overtook on the Artificial Analysis index. It supports text and image inputs with text-only output, a 500,000-token context window, function calling, structured outputs, and reasoning with effort levels from low to the new extra-high setting, letting teams match compute to task difficulty.

The Bigger Picture: Agents Become the Product

The Grok 4.6 release lands one day after SpaceXAI launched Grok Bot, a system that assigns AI agents to run designated tasks as persistent digital coworkers, and weeks after SpaceX's acquisition of Cursor, the AI coding environment where the model debuts. The pattern is deliberate: the frontier race has shifted from single-turn chatbots to long-running, self-verifying agents that complete entire projects, and pricing strategy — making agent-heavy workloads cheap enough to run — is now as competitive as benchmark scores. For developers and enterprises, Grok 4.6 is the strongest sign yet that the most important frontier metrics of 2026 are staying on task, checking your own work, and keeping the inference bill down while doing it.

Sources

whatsapp me