What is the difference between kimi-k2-thinking and regular K2?
It is built on Kimi K2, with a focus on step-by-step reasoning and alternating between reasoning and function calls. When choosing, consider whether the task requires multi-step analysis, tool feedback, and continuous revision, rather than treating Thinking as another name for regular K2. Be sure to explicitly use kimi-k2-thinking when calling it.
Can I use reasoning_effort to adjust thinking intensity?
Do not apply K3's reasoning_effort to this model, and do not apply K2.6's thinking switch either. When using K2 Thinking, organize tasks through clear task boundaries, output requirements, and budget settings; temperature=1.0 is the natively recommended sampling value, not a reasoning intensity level.
How do I read the answer and reasoning content?
In a Chat Completions response, the main text is in message.content, and reasoning content may be in message.reasoning_content. For streaming calls, concatenate the corresponding increments separately to avoid mixing them in the same display area; if tool_calls are returned, handle the tool requests first, then continue retrieving the final answer.
Will functions be executed automatically after tools are integrated?
Tool definitions in Chat Completions tell the model which functions are available; they do not execute functions on behalf of the application. After receiving a call, you need to parse the parameters, execute the corresponding tool, and return the result as a tool message. It is recommended to also handle parameter validation, execution errors, and tool permissions.
How do I call kimi-k2-thinking using the standard API?
Submit model=kimi-k2-thinking and messages to /v1/chat/completions. Read regular results from choices[].message.content; use stream to obtain incremental results for streaming calls. Use this platform's API Key, and set the complete base URL according to the SDK you use.