BATCH TRANSCRIPTION

MAI-Transcribe-2 vs GPT-Transcribe

A product-focused comparison for long podcast and subtitle work: timing, diarization, throughput and the cost of a publishable pass.

The short answer

Choose the contract your workflow can verify.

CastTranscript uses MAI-Transcribe-2 for uploaded media because the current price and transcription-specific output fit a batch publishing workflow. The implementation still probes every provider response: model names do not guarantee that a gateway exposes every timing or speaker field.

Decision pointMAI-Transcribe-2GPT-Transcribe
Product role

Speech recognition for files, subtitles and transcript production

File and Realtime speech-to-text option in the OpenAI model family

CastTranscript role

Default upload and explicit live-refine pass

Comparison benchmark, not a hidden fallback

Published price

$0.10 per audio hour as a limited offer through the end of 2026

Check the current OpenAI pricing page for the selected transcription model

Timing and speakers

Microsoft lists word timestamps, diarization and keyword biasing

Response shape and diarization depend on the selected model and endpoint

Language workflow

Microsoft lists 60 languages plus code switching

Language hints and behavior depend on the chosen transcription model

When fields are missing

Keep verified sentence cues and disclose the capability downgrade

Use the exact API response; do not infer alignment

01

Cost is only the first filter

Storage, chunk retries, summaries and human review all sit around the speech request. A low transcription rate creates room for those parts, but it does not excuse an unverified output contract.

02

Do not stitch models together

Running one model for text and another for speakers or timing can create subtitles that look precise while the words and cues disagree. One speech model owns each pass.

03

Probe the response you receive

A gateway can expose fewer fields than the underlying model. Every job records whether segment timing, word timing and speaker labels were actually returned.

CastTranscript decision

MAI handles batch work. Missing detail stays visible.

Uploads use one MAI pass. If word timing or diarization is absent, the project keeps the trustworthy transcript and sentence cues, disables unsupported exports when necessary, and tells the editor what was not returned.

Try an upload

Questions, answered

MAI and GPT transcription questions

Which model does CastTranscript use for uploaded files?

The default batch pass uses MAI-Transcribe-2. CastTranscript records the provider and model on each job so the model chain is visible.

Does MAI-Transcribe-2 always return word timestamps and speakers?

Microsoft lists word timestamps and diarization as model capabilities, but gateways can expose different response fields. CastTranscript checks the actual response and reports which data were returned.

Will CastTranscript run Whisper or Muse to fill missing fields?

No. The batch path does not silently combine multiple speech models. It delivers verified fields from the MAI pass and shows a capability warning when detail is missing.

Is the $0.10 per hour price permanent?

No permanent price is assumed. Microsoft described $0.10 per audio hour as a limited offer through the end of 2026, so this page includes its verification date.

Private by default

Test the workflow on a recording you know.

Compare the transcript against the source, correct it in one document and export only the formats supported by its timing data.