- SpaceXAI has unveiled Grok 4.7, a larger model designed for coding, knowledge work and complex tasks that can take several hours.
- The release matters because company and independent tests indicate stronger agentic performance, although results vary by runtime environment and workload.
- Grok 4.7 scored 46.3% on CursorBench 4.0, ahead of GPT-5.6 Sol but behind Fable 5.1.
SpaceXAI unveiled Grok 4.7 with a larger base model than Grok 4.6 and a longer reinforcement-learning stage. The developers said the model handles long contexts better, checks its outputs more thoroughly and is optimized for use in Grok Build.
Grok 4.7 is available through Cursor, Grok Build, the API and several third-party platforms. It costs $2 per 1 million input tokens and $6 per 1 million output tokens. A Fast version runs twice as fast and costs twice as much.
Benchmark performance and pricing
In SpaceXAI’s published tests, Grok 4.7 scored 46.3% on CursorBench 4.0, up from 40.4% for Grok 4.6. GPT-5.6 Sol scored 41.7%, while Fable 5.1 led with 51.8%.
On DeepSWE v1.1, Grok 4.7 scored 71%, compared with 72.7% for GPT-5.6 Sol and 70% for Fable 5.1. On terminal-based tasks, Grok 4.7 posted 38%, against 20.3% for Grok 4.6, 37.3% for GPT-5.6 Sol and 57.9% for Fable 5.1.
Grok 4.7 scored 64% on EEBench, which covers electrical-engineering tasks, exceeding both competitors in SpaceXAI’s table. It trailed them on HealthBench Professional.
SpaceXAI listed GPT-5.6 Sol at $4 per 1 million input tokens and $20 per 1 million output tokens, while Fable 5.1 cost $10 and $50, respectively. Grok’s lower price does not translate into leadership across every benchmark, but it strengthens the model’s position in compute-heavy workloads.
Cybersecurity safeguards
SpaceXAI said it fully revamped Grok 4.7’s protective mechanisms. In the company’s HackerBench v0.3 test, the model allowed 3.3% of prompts that SpaceXAI classifies as dangerous dual-use scenarios. The developers said the system also seeks to avoid blocking legitimate work by security professionals.
Some partners received access to Grok 4.7 for closed red-team testing. SpaceXAI published those results itself, meaning they do not constitute an independent assessment.
Independent evaluations show mixed results
Artificial Analysis reported that Grok 4.7 scored 46 points in its Intelligence Index, two points more than Grok 4.6. The platform identified particularly notable progress on long agentic tasks. In its Coding Agent Index, the combination of Grok 4.7 and Grok Build scored 56 points, compared with 47 for the previous version.
The improvement came with higher resource use. In xhigh mode, Grok 4.7 generated about 81,000 output tokens per Intelligence Index task, more than twice the amount used by the previous version in a comparable mode. Artificial Analysis said the main gains were concentrated in agentic work, while changes in some other tests were minor.
Vals AI reported a less favorable outcome. Grok 4.7 scored 54.2% and ranked 24th in its test suite, down from Grok 4.6’s 59.2% score and 14th-place ranking. The largest improvements appeared in legal and medical tasks. The difference from Artificial Analysis’ findings illustrates how evaluations can depend on methodology and the scenarios tested.
Cybersecurity company XBOW also tested Grok 4.7 after receiving early access. In a standard environment, the model slightly underperformed Grok 4.6 in several offensive-security tests and sometimes needed more iterations to achieve the intended result.
Performance improved when XBOW used Grok Build. In one test suite, systems running in that environment increased the number of correctly identified issues from 42 with Grok 4.6 to 68 with Grok 4.7.
XBOW concluded that Grok 4.7’s distinguishing feature was not its ability to solve fundamentally new classes of tasks, but its more efficient performance within suitable agentic infrastructure. The company still ranked its offensive-security quality below the strongest systems it had tested.
Overall, the early independent assessments indicate progress in long-horizon agentic work and programming, while also showing that Grok 4.7’s advantages are not universal and may depend on the toolchain, reasoning level and workload.
Source: Incrypted
