All models

kimi-k2-thinking

MoonshotChatReasoning
Get your API key
kimi-k2-thinking

A thinking model for complex reasoning and multi-step tool collaboration

kimi-k2-thinking is a thinking model built by Moonshot AI based on Kimi K2, focused on tasks that require decomposition, verification, and multi-step progress. It combines reasoning with function calling, making it suitable for code debugging, complex problem analysis, and long-form writing. It natively supports 256K context, can organize materials through text conversations, and can continue completing tasks using execution results after connecting to tools.

MoonshotModel brand
ChatModel type
ReasoningTask capability
STANDARD APIs · QUICK SETUP

Keep your SDK. Connect in minutes.

Point the Base URL to api.acedata.cloud, configure your platform API key and the model ID below, and use your compatible SDK or client.

API hostapi.acedata.cloud
modelkimi-k2-thinking
OpenAI Python SDK
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ACEDATACLOUD_API_KEY"],
    base_url="https://api.acedata.cloud/v1",
)
response = client.chat.completions.create(
    model="kimi-k2-thinking",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

Choose an available protocol for this model. OpenAI SDK uses a Base URL ending in /v1; Anthropic SDK uses the root URL. See each guide for protocol-specific parameters, tools and response formats.

Specifications and API features

Clarify capacity, input and output, and invocation methods before selecting a model.

Native context
256K Token
Native architecture
MoE; 1T total parameters, 32B activated parameters
Native quantization
INT4, using quantization-aware training
Input and output
Text message input; text content, can return reasoning_content
Tool collaboration
Function calling; supports tool definitions and result injection
Response method
Chat Completions supports standard JSON and SSE streaming responses
Sampling recommendation
Native recommendation: temperature=1.0

Architecture and context are native model specifications; this platform provides text chat and session entry points, while tool execution depends on the entry point and application configuration used.

Core Capabilities

Learn what kimi-k2-thinking can bring to your work.

Break complex problems into verifiable steps

K2 Thinking focuses on step-by-step reasoning and multi-step instruction handling. When faced with questions involving multiple conditions, you can ask it to first organize the constraints, then develop a plan and check the conclusion. It is well suited to analytical tasks that require explaining the basis for judgments; prompts should clearly specify the goal, known conditions, and acceptance criteria rather than providing only a broad topic.

Use tool results in subsequent judgment

The model can propose function calls during reasoning and continue analysis based on the returned results. Applications can provide query, calculation, or testing functions to move tasks from textual recommendations to result validation. It is responsible for selecting calls and organizing parameters, while actual execution is completed by the tools, allowing execution permissions to be managed separately from the model's reasoning capabilities.

Analyze and write based on long materials

A native 256K context provides room to process longer texts, code snippets, and ongoing discussions. You can submit background materials together with writing requirements and ask for an outline, argumentation, and a finished draft. When there is substantial material, it is recommended to retain section labels and citation locations so conclusions can be traced back to specific paragraphs rather than merely generating a general summary.

Use Cases

Start from specific tasks to identify where the model can be effective.

Code debugging and repair plans

Provide relevant code, error logs, expected behavior, and the runtime environment, and ask the model to identify possible causes, propose changes, and design tests. After connecting testing tools, you can also feed back failure information for further revision. Deliverables can include patch suggestions, explanations of causes, and a regression test checklist; without execution tools, the output remains a code proposal to be verified.

Multi-criteria solution evaluation

Provide candidate solutions, resource constraints, and evaluation criteria, and let the model analyze trade-offs item by item, identify missing information, and form recommendations. This is suitable for tasks such as technology selection and project planning that should not be decided based on a single metric. You can request a comparison table, key assumptions, and a next-step validation plan to facilitate team discussion rather than directly accepting a conclusion.

Organizing and drafting long materials

Provide multiple text excerpts, meeting notes, or research notes, specify the audience, sections, and argumentation requirements, and generate a structured report draft. You can first organize the relationships among the materials, then expand and revise section by section. Deliverables are suitable for including a summary, main text, and items to be supplemented; real-time information should be obtained through retrieval tools before being incorporated into the analysis.

How to choose this model

Choose based on task complexity, input materials, and expected results.

Choose Thinking when sustained reasoning is needed

If the task involves multiple constraints, repeated verification, or tool feedback, K2 Thinking is a better fit than usage that only requires generating an answer in one pass. It is the model in the K2 series specifically designed for thinking and tool collaboration, and should not be mixed with K2 0905. For brief rewrites or fixed-format filling, first determine whether the task truly requires this kind of reasoning process.

Distinguish Turbo from subsequent series models

kimi-k2-thinking-turbo is an independent public model, not an alias for this model. When switching, use the same set of tasks to compare response experience and result quality without assuming a fixed speed difference. K2.6's thinking switch and K3's reasoning_effort are also not control methods for this model; when migrating, prompts and parameter configurations should be designed separately.

Get started

From a small-scale task to full integration.

01

Prepare tasks and materials

Clarify the objective, required inputs, and output requirements, and use real business examples as a starting point.

02

Try it in the API debugging area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to view the results.

03

Integrate according to the API documentation

Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage boundaries

Understand output quality and capability scope before formal use.

  • A native 256K context does not mean an answer of the same length can be generated in a single request. Long tasks require considering the budget for input, reasoning, and main text together; if a response ends with length, check the output budget or split the task into an outline, section-by-section generation, and final merging.
  • Function calling does not mean the model can run code or access web pages on its own. When using Chat Completions, the application needs to execute tools and return the results. Tool errors, incomplete returned content, or insufficient permissions will all affect subsequent judgment; operations involving writing and publishing should have separate confirmation mechanisms.
  • This model is primarily intended for text reasoning and should not be treated as a native vision or audio model. Materials such as images and scanned documents should preferably be converted into analyzable text first; conclusions from complex reasoning still require checking key assumptions, while code tasks should use actual test results as acceptance criteria.

Frequently Asked Questions

Answers to common questions about using kimi-k2-thinking.

What is the difference between kimi-k2-thinking and regular K2?

It is built on Kimi K2, with a focus on step-by-step reasoning and alternating between reasoning and function calls. When choosing, consider whether the task requires multi-step analysis, tool feedback, and ongoing revision, rather than treating Thinking as another name for regular K2. Be sure to explicitly use kimi-k2-thinking when calling it.

Can I use reasoning_effort to adjust thinking intensity?

Do not apply K3's reasoning_effort to this model, and do not apply K2.6's thinking switch either. When using K2 Thinking, organize tasks through clear task boundaries, output requirements, and budget settings; temperature=1.0 is the natively recommended sampling value, not a reasoning intensity level.

How do I read the answer and reasoning content?

In the Chat Completions response, the main text is in message.content, and the reasoning content may be in message.reasoning_content. For streaming calls, concatenate the corresponding deltas separately and avoid mixing them in the same display area; if tool_calls are returned, handle the tool request first, then continue to retrieve the final answer.

Will functions be executed automatically after connecting tools?

Tool definitions in Chat Completions tell the model which functions are available; they do not execute functions on behalf of the application. After receiving a call, you need to parse the parameters, execute the corresponding tool, and return the result as a tool message. It is recommended to also handle parameter validation, execution errors, and tool permissions.

How should I choose between the two calling endpoints?

If you need to maintain message history yourself, read reasoning content, and control the function-calling flow, choose /kimi/chat/completions. If you want to continue a discussion by session ID, choose /aichat2/conversations, submit question and retain the returned id; continue the conversation later using the same id.

Model resources · Updated: 2026-10-01. For call parameters and billing rules, see the API and pricing sections.