SpaceXAI has launched Grok 4.7, its latest frontier AI model aimed at coding, agentic tasks, and professional knowledge work. The company says the new model is built to handle longer-running tasks, check its own work more carefully, and maintain context across more demanding workflows.
According to SpaceXAI, Grok 4.7 uses a larger base model than Grok 4.6 and was trained through a longer reinforcement-learning process focused on more difficult tasks, particularly problems that can take hours to complete.
The company says the model has improved at verifying its own work and managing longer contexts. It has also been trained to natively understand the Grok Bot harness, which SpaceXAI says improves its performance on conversational tasks and broader knowledge work.
That focus puts Grok 4.7 squarely in the growing category of AI systems designed to do more than answer individual prompts. SpaceXAI is targeting workflows where the model can spend considerably more time working through a problem, particularly software development and professional tasks.
The model also improves on Grok 4.6 across several of SpaceXAI’s reported benchmarks. On CursorBench 4.0, for example, Grok 4.7 scored 46.3%, compared with 40.4% for Grok 4.6. Its Terminal-Bench 4.0 score increased from 20.3% to 38.0%, while its AA Briefcase score rose from 1,546 to 1,657.
SpaceXAI also reports improvements in areas beyond software development. Grok 4.7 scored 64.0% on EEBench for electrical engineering, compared with 53.0% for Grok 4.6, while its Harvey Legal Agent Benchmark score increased from 15.8% to 19.6%. On HealthBench Professional, it reached 56.7%, up from 48.5% for the previous model.
These figures are SpaceXAI’s own benchmark results, so they should be treated as vendor-reported measurements rather than independent evaluations.
SpaceXAI is also highlighting a new safeguard system for Grok 4.7.
The company says the model has been designed to improve both refusal behavior and resistance to jailbreak attempts. In its reported testing, Grok 4.7 achieved a 62.4% score on LatchBio’s biosafety benchmark.
For cybersecurity, SpaceXAI says Grok 4.7 allows only 3.3% of risky dual-use prompts through on its HackerBench v0.3 evaluation, while maintaining low refusal rates for legitimate security work. The company has also begun providing select cybersecurity partners with invite-only access to the model’s red-team capabilities for defensive research.
One of the more notable parts of the launch is that SpaceXAI has kept Grok 4.7’s standard API pricing at the same level as Grok 4.6.
The model starts at $2 per million input tokens and $6 per million output tokens. Cached input is priced at $0.50 per million tokens, while requests exceeding 200,000 prompt tokens are charged at higher rates. Grok 4.7 supports a 500,000-token context window and reasoning levels ranging from low through xhigh.
SpaceXAI is also offering a faster variant. Grok 4.7 Fast uses the same underlying model with twice the output speed, but costs twice the standard token rates. The Fast version is currently available through Cursor and Grok Build rather than the public xAI API.
Grok 4.7 is available now through the SpaceXAI API, Grok Build, Cursor, third-party coding harnesses, model routers, and cloud platforms. The company is also making it available in Grok Build’s free tier.
With Grok 4.7, SpaceXAI is continuing to push Grok toward a role as an agent that can work through lengthy coding and professional tasks rather than simply responding to individual questions. The combination of a larger model, longer reinforcement learning, stronger self-verification, and unchanged base API pricing makes this a substantial upgrade over Grok 4.6 on paper.
