What is the difference between kimi-k2-thinking and regular K2?
It is built on Kimi K2, with a focus on step-by-step reasoning and alternating between reasoning and function calls. When choosing, consider whether the task requires multi-step analysis, tool feedback, and ongoing revision, rather than treating Thinking as another name for regular K2. Be sure to explicitly use kimi-k2-thinking when calling it.
Can I use reasoning_effort to adjust thinking intensity?
Do not apply K3's reasoning_effort to this model, and do not apply K2.6's thinking switch either. When using K2 Thinking, organize tasks through clear task boundaries, output requirements, and budget settings; temperature=1.0 is the natively recommended sampling value, not a reasoning intensity level.
How do I read the answer and reasoning content?
In the Chat Completions response, the main text is in message.content, and the reasoning content may be in message.reasoning_content. For streaming calls, concatenate the corresponding deltas separately and avoid mixing them in the same display area; if tool_calls are returned, handle the tool request first, then continue to retrieve the final answer.
Will functions be executed automatically after connecting tools?
Tool definitions in Chat Completions tell the model which functions are available; they do not execute functions on behalf of the application. After receiving a call, you need to parse the parameters, execute the corresponding tool, and return the result as a tool message. It is recommended to also handle parameter validation, execution errors, and tool permissions.
How should I choose between the two calling endpoints?
If you need to maintain message history yourself, read reasoning content, and control the function-calling flow, choose /kimi/chat/completions. If you want to continue a discussion by session ID, choose /aichat2/conversations, submit question and retain the returned id; continue the conversation later using the same id.