Can GPT-4.1 analyze images directly?
It can understand text and images. When using Chat Completions, combine text and image_url content blocks in the same message, and specify the chart, area, or issue to inspect. Results are returned as text analysis; if you need to generate or edit images, choose a dedicated image model.
Does 1 million context mean it can generate answers of the same length?
No. Context capacity and the single-output limit are different metrics: GPT-4.1 has a native context window of up to 1 million tokens and a maximum output of 32,768 tokens. Long-form deliverables can be generated by chapter, while reserving space for input materials, conversation history, and responses.
Should I use Responses or Chat Completions?
Applications with existing message arrays can use Chat Completions, submitting model and messages and reading the response in choices. Responses uses model and input, processing output by response events. The data structures for the two entry points differ, so request bodies or parsing logic should not be mixed directly.
How does GPT-4.1 maintain multi-turn conversations?
When using a message-based interface directly, you need to provide relevant history with each request. To reduce client-side history management, you can choose the AI Chat session entry point, set stateful, and include the returned id in subsequent requests. Saved sessions help continue conversations, but do not mean that all history is retained completely forever.
How can I make GPT-4.1 output processable JSON?
Chat Completions provides json_object and json_schema format settings. It is recommended to clearly specify field meanings, required fields, and rules for missing values, and to validate both structure and business logic on the receiving end. Valid JSON only means the format can be parsed; it does not guarantee that amounts, categories, or cited content are correct.