Kimi K2.7 Code: 3 Design Drafts in 8 Minutes, 3 Bugs Fixed in 20

Β·Toolin Editorial Team

Kimi K2.7 Code supports parallel agents (swarm), forces thinking on, and its coding ability approaches GPT-5.5 and Opus 4.8

Kimi K2.7 Code: 3 Design Drafts in 8 Minutes, 3 Bugs Fixed in 20

Kimi K2.7 Code is the new-generation coding model released and open-sourced by Moonshot AI, now switched on across Kimi Code in full. Its defining traits are speed, low cost, and support for parallel agents (swarm), with some coding benchmarks approaching GPT-5.5 and Opus 4.8. A high-speed version launches next Monday, raising output speed 5-6x.

Core Features

Parallel Agents (Swarm)

K2.7 Code's standout ability is proactively spinning up parallel agents. Give it a design task and it automatically launches multiple agents working at the same time, each producing a complete draft in a different style.

Hands-on example: given a single design Skill instruction, K2.7 Code launched three parallel agents within 8 minutes, each producing a complete HTML animation mock in a different design school, and automatically took screenshots for selection.

Forced Thinking Mode

The model forces thinking on (similar to Claude Fable 5's approach); turning it off makes the API throw an error outright. The team specifically tuned the overthinking problem, using 30% fewer thinking tokens than the previous generation. The thinking process is displayed throughout, which greatly reduces the wait-time anxiety.

Automated Debugging

In testing, K2.7 Code was asked to fix three bugs in FanBox (a product built with Claude Fable 5). The full rundown:

  • When special bytes made the file-reading tool refuse to read a file, it switched to Python binary inspection on its own
  • When a local proxy hijacking localhost caused a 502, it investigated processes and environment variables by itself and added parameters to route around it
  • After fixing, it automatically ran syntax checks and built three edge-case test pages to verify the results

Three bugs, fixed in 20 minutes, without a single line of code typed by hand.

Performance and Pricing

ItemStandardHigh-speed (launches next Monday)
Output speedStandard180 token/s (260 on short context)
API input price6.5 CNY per million tokens-
API output price27 CNY per million tokens-
Cache hit1.3 CNY per million tokens-
Price multiplier1x2x (6x speed)

Once the high-speed version launches, the 8-minute animation design can shrink to about 2 minutes, and the 20-minute bug fix to about 4 minutes.

How to Use It

Via the Kimi Code CLI

  1. Visit kimi.com/code
  2. Membership starts at 49 CNY, with quota counted on a rolling 5-hour window
  3. Supports custom Skills and can reuse an existing skills directory directly

Via the API

  • Input: 6.5 CNY per million tokens
  • Output: 27 CNY per million tokens
  • Cache hit: 1.3 CNY per million tokens
  • High-speed version: 2x the price, 6x the speed

Use Cases

  • Rapid prototype design: parallel agents produce multiple versions at once
  • Bug fixing and maintenance: a closed loop of independent investigation, fixing, and verification
  • HTML animations / interactive pages: design expressiveness beyond expectations
  • Skill-driven workflows: design the loop once, agents run it over and over

Suggested Division of Labor

ModelProfileBest for
Claude Fable 5 / Opus 4.8Slow, precise, expensiveBreaking new ground from zero to one, design loops
Kimi K2.7 CodeFast, steady, cheapRunning nonstop inside the loop, day and night
GLM-5.2Open source, locally deployableData-sensitive scenarios, the top pick in China

πŸ’‘ Tip: Kimi Code recognizes Claude Code's skills directory directly, no reinstall needed. If you can't remember a Skill's name, type /skill and pick from the dropdown list.


Sources: hands-on testing by Huashu, Kimi official announcements