Grok 4.5: xAI's Coding Model Co-Trained with Cursor, Starting at $2
xAI releases Grok 4.5, co-trained with Cursor and built for real engineering tasks. API pricing is $2/$6 per million tokens, and it's live on all Cursor plans.


Grok 4.5: xAI's Coding Model Co-Trained with Cursor, Starting at $2
xAI releases Grok 4.5, co-trained with Cursor and built for real engineering tasks. API pricing is $2/$6 per million tokens, and it's live on all Cursor plans.
On July 8, xAI released Grok 4.5 β a coding model co-trained with Cursor. The story here isn't "yet another general-purpose frontier model"; it's that xAI fed real coding data from Cursor's developer base into the training set, with a very clear goal: build a coding model aimed at real engineering tasks and cheap enough to deploy at scale.
Core Positioning
Grok 4.5 is a mixture-of-experts (MoE) model aimed at coding, agent tasks, and knowledge work. xAI's release notes stress two points:
- Trained on tens of thousands of NVIDIA GB300 GPUs
- The training data includes trillions of tokens of real Cursor coding data
That means the model wasn't trained on synthetic data or contest problems β it's aligned directly with developers' real workflows inside the IDE.
Benchmarks on Real Engineering Tasks
Across several engineering benchmarks, Grok 4.5 performs as follows (official xAI figures):
| Benchmark | Grok 4.5 | Comparison (Fable/Claude Opus 4.8/GPT-5.5) |
|---|---|---|
| SWE Marathon (pass@1) | 29.0% | Opus 4.8 26.0% / Fable 24.0% |
| Terminal Bench 2.1 | 83.3% | Fable 84.3% / GPT-5.5 83.4% |
| SWE Bench Pro | 64.7% | Fable 80.4% / Opus 4.8 69.2% |
| DeepSWE 1.0 | 62.0% | Fable 66.1% / GPT-5.5 64.31% |
Worth noting: on the most hardcore engineering benchmarks like SWE Bench Pro and DeepSWE, Grok 4.5 still trails Fable 5 (Anthropic). But on long-horizon tasks like SWE Marathon, it beats both Opus 4.8 and Fable.
π‘ How to read these numbers: if you purely want the strongest coding capability, Fable 5 / the Claude line still lead. Grok 4.5's edge is value for money β the same budget runs more tasks.
Token Efficiency: 4.2x
This is Grok 4.5's hardest cost advantage. On SWE Bench Pro tasks:
- Grok 4.5: an average of 15,954 output tokens
- Opus 4.8 (max): an average of 67,020 output tokens
In other words, Grok 4.5 solves the same tasks with about 1/4.2 the tokens of Opus 4.8. Layer on the lower unit price and the real cost gap widens further.
Pricing
Grok 4.5's API pricing:
- Input: $2 per million tokens
- Output: $6 per million tokens
Model speed is 80 TPS (tokens per second) β the "fast model" tier.
A third-party cost analysis (the-decoder) ran the numbers on "cost per solved task":
- Grok 4.5 (Grok Build): $2.49 per task
- GPT-5.5 (Codex): $5.07 per task
- Fable 5 (Claude Code): $11.80 per task
Grok 4.5 comes in at roughly half of GPT-5.5 and a fifth of Fable 5.
How to Use It
Grok 4.5 is already available through:
- Cursor: usable on all plans (the core beneficiary of the xAI-Cursor co-training)
- Grok Build: Grok 4.5 is the default model, free for a limited time
- SpaceXAI console: grab an API key and call it directly
# Call via the SpaceXAI API (illustrative; see the x.ai API docs for details)
# 1. Create an API key in the x.ai console
# 2. Call Grok 4.5 over the OpenAI-compatible protocolβ οΈ Regional restriction: as of release, Grok 4.5 is not yet live in the EU region (including SpaceXAI products and the API console), with availability expected in mid-July.
Office Capabilities
Beyond coding, Grok 4.5 in Grok Build also handles Office-style tasks:
- Excel: builds complex financial/data models, uses formulas across sheets, leaves sticky-note comments
- PowerPoint: builds complex charts from native shapes, designs intuitive slides
- Word: writes clear long-form documents
Official plugins are available for Word, PowerPoint, and Excel.
Who It's For
- Cursor users: free on all plans, near-zero switching cost
- Teams deploying coding agents at scale: at $2.49 per task, well suited to running large volumes of SWE-style work
- Budget-sensitive developers: 4.2x token efficiency plus a low price suits high-frequency calling
- Users chasing peak coding capability: for hardcore tasks like SWE Bench Pro, Fable 5 / Claude is still the recommendation