Is gpt-transcribe the same as gpt-4o-transcribe?
When using it, treat gpt-transcribe as a standalone invocation ID. Do not replace it with gpt-4o-transcribe or gpt-4o-mini-transcribe on your own. Similar names do not mean identical versions; when choosing, focus more on your own recording samples and required delivery format.
How do I submit audio when calling it directly?
Submit a binary file to /v1/audio/transcriptions, and explicitly specify model=gpt-transcribe. Do not rely on the default selection when the model is omitted. If submitting a URL through an MCP audio transcription tool, follow that tool's input method; do not treat the URL string directly as an uploaded file.
Can it generate SRT or VTT subtitles directly?
The endpoint provides srt and vtt request options. You can first test a short recording to see whether it produces subtitle results that meet your needs. Before final delivery, also check sentence segmentation, time alignment, and proper nouns; if your workflow mainly uses plain text, you can transcribe first and then arrange the timeline during subtitle editing.
How should I prepare a request when there is a lot of specialized vocabulary?
You can prepare concise language information and terminology context around language, prompt, or keywords[], and first test common names and easily confused terms. Prompts are not dictionaries that guarantee correct recognition, nor are they suitable for including rewriting instructions; key terms should still be checked item by item against the original recording.
Can I get real-time responses while speaking?
gpt-transcribe's primary workflow is audio-file transcription, and its output is recognized text rather than conversational responses. Even when using streaming configuration, it should not be considered equivalent to a real-time two-way voice conversation; to listen and respond while speaking, you also need to design separate audio capture, conversation processing, and speech playback stages.