Kimi K2.7 Code Released: Token Consumption Down 30%
Moonshot AI has released and open-sourced the Kimi K2.7 Code coding model: 1.1 trillion parameters, 256K context, a big fix for overthinking on long-horizon tasks, plus a Speed variant with 6x the speed at 2x the price.


Kimi K2.7 Code Released: Token Consumption Down 30%
Moonshot AI has released and open-sourced the Kimi K2.7 Code coding model: 1.1 trillion parameters, 256K context, a big fix for overthinking on long-horizon tasks, plus a Speed variant with 6x the speed at 2x the price.
Moonshot AI today released and open-sourced the Kimi K2.7 Code coding model. 1.1 trillion parameters, a 256K context window, and a fix for a long-standing pain point: the "overthinking" problem in long-horizon coding tasks. Average token consumption is down 30%, and a Speed variant launches next Monday with 5-6x the speed at only 2x the price.
Key Upgrades
Compared with the previous-generation K2.6, K2.7 Code shows clear gains on three dimensions:
Improved coding ability
- Kimi Code Bench v2 up 21.8%
- Program-Bench up 11%
- MLS Bench Lite up 31.5%
Improved agent ability
- Roughly 10% performance gains on the Kimi Claw 24/7 Bench, MCP Atlas, and MCP Mark Verified benchmarks
Improved token efficiency
- Average token consumption down 30%
- Higher performance with fewer tokens
Pricing and Availability
API pricing (same as K2.6):
- Standard input: ¥6.5 per million tokens
- Standard output: ¥27 per million tokens
- Cached input: ¥1.3 per million tokens
How to get it:
- API: via the Kimi API open platform
- Coding plan: the default model on kimi.com/code has been upgraded in sync
- Open source: model weights are available for download on Hugging Face
Note: Thinking mode must be enabled to use K2.7 Code. Turning it off manually causes API errors, and Kimi Code falls back to the K2.6 model. Both the API and Kimi Code enable it by default.
Speed Variant: 6x the Speed, 2x the Price
The K2.7 Code Speed variant launches next Monday (June 15):
- Output speed of about 180 Token/s in typical coding scenarios
- Up to 260 Token/s for short-context scenarios
- Priced at 2x the standard version
Through the Kimi Code Plan early access program, you can try the Speed variant inside Kimi Code, with gradual rollout to Allegretto members and above expected in July.
Limited-Time Promotion
To celebrate the K2.7 Code launch, the Kimi API open platform is running a three-week top-up bonus: top up ¥500 or more and receive a 20%-30% voucher bonus. Vouchers are consumed first when used and are valid for 90 days.

Hands-On Experience
In hands-on testing, the most noticeable impression of K2.7 Code is that it has become "more decisive." The old habit of second-guessing itself on simple tasks and deliberating at length before acting has dropped off sharply.
Test case 1: macOS-style frontend demo
Recreate a macOS-style operating system interface in a single HTML file. K2.7 Code didn't spend much time thinking and jumped straight into development, ultimately delivering a complete boot animation and working basics (Notes, the browser, and other apps all function properly).

Test case 2: agent town
We had K2.7 Code build a recreation of Stanford's "agent town." It first generated a PRD covering product overview, market context, functional architecture, and the technical approach, then developed the MVP under the PRD's guidance. After 30 minutes of continuous iteration it delivered a fully working project with a clean file structure and sensible division of labor.
Quick Start
Coding with K2.7 Code:
- Kimi Code monthly plan: kimi.com/code
- API quick start: platform.kimi.com
- Local deployment: head to Hugging Face to download the model weights
Best use: K2.7 Code is recommended for coding tasks; for non-coding tasks, the more well-rounded K2.6 remains the pick.
Related articles

Xiaomi MiMo UltraSpeed: A Trillion-Parameter Model Running at 1000 tokens/s
Xiaomi's MiMo-V2.5-Pro UltraSpeed delivers 1000 tokens/s on a trillion-parameter model running on generic 8-GPU hardware — and it's the full-strength Pro version, not a dumbed-down Flash variant. The API is live and taking applications.

Tabbit: A Permanently Free AI Browser with 10+ Top Models at Your Fingertips
Meituan launches AI browser Tabbit V1.0 with core features permanently free, 10+ top domestic large models and agent capabilities built in, plus one-click access to 300+ ready-made trick skills.

Finding Overseas Influencers with AI Employees: A Hands-On AhaCreator Guide
A step-by-step walkthrough of the full overseas influencer marketing workflow with AhaCreator — from creator sourcing and content review to cross-border payouts. Ideal for indie developers and teams going global.

AI Long-Video Generation: A Head-to-Head Review of Two Open-Source Frameworks
VideoClaw and JoyAI-Echo, two open-source frameworks released the same day, tackle AI long-video consistency through multi-agent collaboration and cross-modal memory banks respectively — this article compares their technical approaches.

ChatGPT's Memory System Gets a Full Overhaul: Dreaming V3 Is Live
OpenAI ships the new Dreaming V3 memory architecture — ChatGPT now "dreams" in the background to organize what it knows about you. Free access for 1 billion users for the first time, with doubled memory capacity for Plus/Pro.

Cloudflare Integrates Claude Managed Agents: A Developer's Practical Guide
Cloudflare adds support for Claude Managed Agents, letting developers run Claude agents on Cloudflare's platform, connect to private systems, and deploy AI agents securely.