
Grok 4.7 Released: xAI’s Faster AI Model for Coding & Agents
On September 21, 2026, SpaceXAI officially announced the release of Grok 4.7, its latest artificial intelligence model designed for software engineering, reasoning, and long-horizon knowledge work. According to the company’s launch release, the new model delivers up to double the speed at roughly half the price of comparable frontier models, while maintaining the baseline pricing structure and response speed of its predecessor, Grok 4.6.

The upgraded system relies on an extended reinforcement learning training regimen heavily weighted toward demanding programming and mathematical problems that take hours to complete. SpaceXAI notes that Grok 4.7 demonstrates substantial improvements in self-verification, error detection, and long-context management, alongside an overhauled safety stack designed to resist adversarial jailbreaks.
While Elon Musk previously indicated in public statements that Grok 4.7 scales to roughly 2.1 trillion parameters-up from approximately 1.5 trillion in Grok 4.6-SpaceXAI’s official release notes refrained from citing a specific parameter figure, stating instead that the model is powered by a new, larger base model.
Architectural Scaling and Long-Horizon Reinforcement Learning
In its technical overview, SpaceXAI highlighted that Grok 4.7 was trained using an intensive reinforcement learning pipeline focused on complex tasks. Rather than optimizing strictly for rapid conversational answers, engineering teams calibrated the model to handle sustained reasoning across complex, multi-step tasks.
These improvements are designed to help the model verify its work, catch errors during code generation, and maintain coherence across complex multi-step tasks.
While SpaceXAI did not specify the context-window size in its launch announcement, Artificial Analysis currently lists Grok 4.7 with a 500,000-token context window.
Native Integration with Grok Bot and Autonomous Agent Workflows
Beyond standard text prompting, Grok 4.7 was natively fine-tuned to interface with the Grok Bot harness-the autonomous agent environment introduced by SpaceXAI on August 11, 2026.
The Grok Bot framework operates as an always-on digital workforce. Equipped with dedicated cloud computing environments, these autonomous agents interact directly with tools and apps to execute multi-hour tasks around the clock without requiring constant human supervision.
This makes Grok 4.7 particularly suited to autonomous tool use, multi-step execution, and knowledge-work tasks within the Grok Bot environment.
Benchmark Results: Software Engineering and CursorBench 4.0
SpaceXAI published evaluation data comparing Grok 4.7 against leading models across code refactoring, terminal execution, and technical engineering suites:
- CursorBench 4.0: In tests evaluating long-horizon, multi-file code editing, Grok 4.7 scored 46.3% under its extra-high effort setting, outperforming Grok 4.6 (40.4%) and GPT-5.6 Sol (41.7%), while trailing Fable 5.1 (51.8%).
- DeepSWE v1.1: In real-world software engineering benchmarks, Grok 4.7 achieved 71.0% at high effort, surpassing Fable 5.1 (70.0%) and Grok 4.6 (65.2%), while closing the gap with GPT-5.6 Sol (72.7%).
- EEBench (Electrical Engineering): Grok 4.7 led all competing models with a score of 64.0%, compared to 53.0% for Grok 4.6, 56.4% for Fable 5.1, and 39.4% for GPT-5.6 Sol.
- Terminal-Bench 4.0: On multi-hour command-line tasks, Grok 4.7 reached 38.0%, marking an 87% relative improvement over Grok 4.6 (20.3%).
Artificial Analysis Evaluations and Professional Workloads
Evaluations published by Artificial Analysis further assess Grok 4.7’s performance across specialized professional domains and complex office workflows:

- Artificial Analysis Intelligence Index: Grok 4.7 currently scores 46 on Artificial Analysis’ composite Intelligence Index, which combines multiple evaluations covering areas such as coding, knowledge work, agentic tasks, and long-context reasoning.
- AA Briefcase v1.1: Measuring multi-hour office workflows and knowledge work, Grok 4.7 scored 1,657 points, topping Grok 4.6 (1,546) and GPT-5.6 Sol (1,487), and finishing narrowly behind Fable 5.1 (1,678).
- GDPval Elo Rating: In evaluations simulating professional tasks performed by financial analysts, attorneys, and consultants, Grok 4.7 recorded an Elo score of 1,695, surpassing Grok 4.6 (1,605) and GPT-6 Astra (1,542), while Fable 5.1 scored 1,735.
- Harvey Legal Agent Benchmark: Grok 4.7 led the legal category with a score of 19.6%, ahead of Grok 4.6 (15.8%), Fable 5.1 (6.7%), and GPT-5.6 Sol (2.5%).
- HealthBench Professional: In clinical reasoning and diagnostic evaluation, Grok 4.7 scored 56.7%.
Rebuilt Safeguards, Biosecurity, and Cyber Defense
SpaceXAI emphasized that Grok 4.7 features a newly architected safety stack designed to minimize refusal rates for legitimate use while maximizing resistance against jailbreaks and dangerous dual-use queries.
In biological safety evaluations on the LatchBio benchmark, Grok 4.7 achieved a top score of 62.4%, refusing hazardous biological queries without restricting benign scientific research. This builds upon biosecurity monitoring established by Grok 4.6 earlier in September 2026.
On cybersecurity evaluations, the model allowed only 3.3% of risky prompts to pass on HackerBench v0.3, while rarely blocking legitimate security work. SpaceXAI also announced an invite-only red-team research initiative, granting selected cybersecurity partners access to Grok 4.7 for defensive research.
Broad IDE Support and GitHub Copilot Integration
SpaceXAI expanded the distribution of Grok 4.7 across developer tools and platforms from launch.
On September 21, 2026, GitHub announced official support for Grok 4.7 within GitHub Copilot across plans including Copilot Pro, Pro+, Max, Business, and Enterprise tiers. Integration spans major development environments such as Visual Studio Code, Visual Studio, Xcode, JetBrains IDEs, Eclipse, Copilot CLI, and the Copilot cloud agent.
Simultaneously, SpaceXAI made Grok 4.7 accessible within the Cursor editor, via the terminal-based Grok Build agent, through the official Grok API, and across leading cloud providers and model routers. Free evaluation access is also provided inside Grok Build at x.ai/build.
Pricing and Cost Efficiency
Despite architectural updates, SpaceXAI kept standard pricing at $2 per million input tokens and $6 per million output tokens-the same price as Grok 4.6.
The company also offers Grok 4.7 Fast, an accelerated variant delivering twice the output speed at double the price ($4 input / $12 output per million tokens).
By comparison, SpaceXAI’s benchmark tables list GPT-5.6 Sol at $4 per million input tokens and $20 per million output tokens, and Fable 5.1 at $10 input and $50 output per million tokens. For development teams and enterprises processing millions of tokens daily, this pricing structure represents a notable reduction in operational costs.
The Future of Agentic Software Engineering
The launch of Grok 4.7 reflects xAI’s continued focus on coding, agentic workflows, and knowledge work. By combining extended reinforcement learning, competitive API pricing, and integrations across development tools and autonomous agents, the model expands xAI’s push beyond conversational AI toward longer-running software and knowledge-work tasks.




