xAI positions Grok 4.7 for coding and knowledge work, emphasizing longer runs on difficult tasks and more careful self-checking. The company says it uses a larger base model, longer reinforcement-learning training weighted toward multi-hour problems, and improved context management. It is available through the Grok API, Grok Build, and Cursor; standard API input and output prices match Grok 4.6. See xAI’s announcement and evaluation table.

How to read the evaluations
xAI reports results on CursorBench, DeepSWE, Terminal-Bench, GDPval, and AA Briefcase, covering coding, terminal work, and professional office tasks. The announcement compares models and reasoning settings across task types, so one score cannot summarize overall quality. Long-horizon tasks differ from a standard single response, and a high-effort score is not the same as default behavior. These vendor-reported results are useful for choosing what to test, not a guarantee for every project.
Price, speed, and access
Standard API rates are $2 per million input tokens and $6 per million output tokens, unchanged from the previous generation. xAI also offers Fast mode, which the announcement describes as about twice as fast at about twice the price. Whether that is worthwhile depends on wait time, retries, output size, and success rate; token price or speed alone does not tell you the cost of a completed task.
Compare it on your codebase
- Prepare real bug fixes, test additions, and multi-file feature tasks with clear acceptance criteria.
- Run Grok 4.6 and 4.7 with identical repository context, tools, permissions, and prompts.
- Run your automated tests and have developers review the diffs blind. Track first-pass success, error types, human edits, and total token cost.
- For long tasks, ask the model to state a plan and validation steps first. Allow writes only in an isolated branch.
Where it may help
Long-horizon performance can matter more than short-answer leaderboard scores when a task spans files, documents, or repeated tool calls. For simple questions, a more expensive Fast mode may not pay off. Legal, medical, and security material still needs qualified review; benchmark results do not constitute professional advice or replace operational checks.