Model Configuration | Kimi Code Docs

4 min read Original article ↗

Model Configuration

This page covers the models Kimi Code provides and how to switch between them in each client.

Model Overview

Kimi Code currently offers two models—Kimi K3 and Kimi K2.7 Code—across three model IDs, selectable by model ID in clients or third-party tools. Model specs:

Model IDk3kimi-for-codingkimi-for-coding-highspeed
Model versionKimi K3Kimi K2.7 CodeKimi K2.7 Code
DescriptionKimi's most capable flagship coding model: 2.8T parameters, the first open-source model at the 3-trillion-parameter scale; up to 1M context window with low / high / max reasoning levels, excels at complex engineering tasks and long-horizon reasoningA mature, stable coding model that reliably follows instructions with Thinking on and delivers high coding task success ratesThe high-speed version of K2.7 Code, with the same coding ability and ~5–6× faster output
SpeedRegularRegularHighSpeed (6× speed, 3× quota usage)
Context windowUp to 1M256k256k
Reasoningreasoning_effort:low / high / max (default high)Thinking:ONThinking:ON
AvailabilityAndante: not supported
Moderato: 256k context
Allegretto and above: up to 1M context
All membersAllegretto plan or above

Need a higher membership plan?

Different membership plans unlock different models, context windows, and speeds. Upgrade your plan →

Why did usage go up after the new model launched?

After switching models, the context cache built earlier no longer hits on the new model, so that context has to be re-prefilled. Usage therefore looks higher right after switching. Recommended action:

  • Start a new session when using the new model: this gives better results and lower consumption.
Why do I still get a 401 with the correct model ID?

When the requested capability exceeds your plan's entitlements, the server returns 401. Three common cases:

  • No K3 access: your plan is below Moderato and can't call k3 — upgrade to Moderato or above.
  • No 1M access: on a Moderato plan, k3 supports up to 256K context; up to 1M context is available on Allegretto and higher tiers.
  • No HighSpeed access: some plans don't include HighSpeed — upgrade to Allegretto or a higher tier to call kimi-for-coding-highspeed.

For the full error text and how to handle it, see the Error Reference.

Why isn't HighSpeed noticeably faster?

Two common reasons:

  • Mistyped model ID: the HighSpeed ID must be kimi-for-coding-highspeed; a wrong value silently falls back to the standard kimi-for-coding — no error, no speedup.
  • Tools and scripts dominate: HighSpeed only speeds up model output. Tool calls (reading/writing files, running commands, etc.) and script execution are unaffected, so when they take up most of a turn the overall speedup feels small.
How to reduce the overhead of switching reasoning effort?

Switching reasoning effort invalidates the context cache you've built up, so context that would have hit the cache must be re-prefilled. To avoid triggering re-prefill too often:

  • Pick an effort that fits the task and keep it consistent within a session;
  • When you genuinely need a different effort, start a new session rather than switching back and forth in a long session.

How to Switch Models

Usage notes

  • Try K3 in a new session: switching models invalidates the cache you've built up. Start a new session to avoid extra token costs and get a better experience.
  • Fill in the Model ID, not the model version name: when calling a model, use one of the Model IDs from the table above (k3, kimi-for-coding, kimi-for-coding-highspeed). Entering a model version name like Kimi K3 or K2.7 Code will cause the call to fail.
  • K3 / K2.7 without Thinking routes to K2.6: keep Thinking on to use K3 or K2.7 Code; disabling thinking routes the request to K2.6.

Ways to switch to the target model:

Official clients

  • Official Kimi Code CLI: type /model to switch models—no config changes needed; if the latest model isn't listed yet, /logout and sign in again with /login.
  • Kimi Code for VS Code: pick the target model from the dropdown menu in the input bar; if it isn't listed yet, restart VS Code or reinstall the extension.

Third-party tools

Set the tool's Model ID to the target model. Detailed steps:

  1. Create an API Key in the Kimi Code Console.
  2. Fill in the Base URL and the corresponding Model ID in your tool.

Kimi Code API supports both OpenAI and Anthropic protocols. Base URLs:

ProtocolBase URL
OpenAI compatiblehttps://api.kimi.com/coding/v1
Anthropic compatiblehttps://api.kimi.com/coding/

For detailed setup steps, see the corresponding tool guide:

Before using K3 in third-party tools

K3's setup differs slightly from K2.7 Code. Before using it, check the two points below:

  • Context window: some tools default to a context window smaller than the max 1M—manually set the context-window field to 1048576 to use K3's full up-to-1M context.
  • Reasoning effort: K3 supports low / high / max; the effort a tool sends is mapped as below:
# default
null / undefined       → high
any other unknown       → HTTP 400 error

# → max
ultra / max / xhigh     → max

# → high (recommended)
high / medium           → high

# → low
low / minimum / light   → low

# → thinking disabled
none                    → thinking.type disabled