Prepare a source you can use
Start from the original podcast, interview or finished video file. Use clean source audio when possible; music beds, remote-call echo and aggressive noise reduction can all make names and speaker changes harder to recognize.
For YouTube, import captions you can already access or upload an audio track you have permission to process. A public URL is not permission to download the underlying media.
Best starting point A final edit with stable timing. If the picture changes later, regenerate the subtitle file from the new cut.
Generate a transcript with real cue timing
CastTranscript prepares long media as short queued sections, then combines the results into one document. This keeps a multi-hour episode from depending on one long-running background task.
The editor distinguishes sentence timing from word timing. If a provider returns sentence cues but no word-level offsets or speaker labels, the project reports that limitation instead of fabricating alignment with a second model.
Edit for reading, not just recognition
Correct guest names, brands and technical terms first. Then play each uncertain cue, remove verbal debris only when it improves readability, and keep the speaker's meaning intact.
Choose the file your destination expects
SRT is the safest default for video editors and most publishing platforms. WebVTT is designed for web video and supports browser-friendly cue syntax. TXT is best for reading and search; Markdown keeps headings and show notes useful in a CMS.
Run a three-point playback check
Open the exported subtitle file in the destination player and inspect the first minute, a dense exchange in the middle and the final minute. That catches offset errors, overlaps and end-of-file truncation without replaying the entire episode.
Create your transcript