Kimi Code says K2.8 Preview has fully rolled out. The existing kimi-for-coding model ID is upgraded in place, so most clients need no configuration change. Users can select low, high, or max reasoning effort, and eligible plans offer up to a one-million-token context window. Kimi also says coding usage depends on membership; the Go plan does not include coding quota. See the Kimi Code model configuration guide.
What the in-place upgrade means
Keeping the same model ID avoids changing endpoints across clients, but existing sessions, prompts, and cached context may not behave exactly as before. Kimi’s documentation says context cache cannot be reused across different models; starting a fresh session after a switch can avoid extra prefill usage. An unchanged ID does not guarantee identical output style, reasoning time, or quota consumption.
Choose a reasoning level
- Low: Start here for clear, short tasks such as completion, formatting, or a localized fix.
- High: Use for routine development that needs multi-file inspection, an explanation, or test requirements.
- Max: Reserve it for architecture analysis, multi-step debugging, or extended coding. It may take more time and quota, so it need not be the default for every request.
Before a new task, record the selected model, reasoning level, context window, and plan eligibility. Kimi lists different capabilities by plan; a one-million-token context is not automatically available to every account.
A short migration check
- Confirm that the client still uses
kimi-for-codingand is not pinning an older version independently. - Use a fresh session to rerun representative tasks: a simple completion, a multi-file change, and a failing-test investigation.
- Compare completion quality, tool behavior, latency, and quota. Record the first run separately because a model switch can affect context-cache usage.
- Document the model ID, reasoning level, and client version so teammates can reproduce the environment.