Interview Transcription: From Recording to Clean Copy
An hour of interview audio takes an afternoon to type by hand and about twenty minutes to proofread with speech recognition - if the pipeline is right: prep, transcribe, verify.
The pipeline: prep, transcribe, verify
Prep first: extract the audio if the interview lives in a video file, then cut dead air and off-topic chatter - recognition time is wasted on silence and proofreading load shrinks with it. Normalize to standard PCM WAV or a high-quality MP3 for the most stable results.
Transcribe locally to get the draft, then verify: read the draft through, mark the rough spots, and re-listen only there - proper nouns, numbers, overlapping speakers. Everything stays on your device, which matters for unpublished reporting.
Audio prep that lifts accuracy
Audio quality caps accuracy. Denoise hissy recordings - a better signal-to-noise ratio visibly cuts proper-noun errors. Overlapping speakers are the worst case: asking interviewees to pause before responding beats any post-processing.
Even out volume: quiet passages drop words. Compression or normalization flattens the dynamics before recognition. Split long recordings into 10-20 minute segments by topic so transcription and proofreading progress in manageable chunks.
Proofreading discipline
Do not re-listen to everything - read the draft, flag the sentences that feel off, and check only those against the recording. Keep a proper-noun table (names, organizations, terms) and run a find-replace pass; it beats fixing each instance.
Decide the output standard up front: verbatim (with fillers, for legal or research use) or edited copy (redundancy trimmed, for media use). Keep section-level timestamps - they anchor every later quote and lookup.