Upload an interview – audio or video – and get it back as a dialogue with the speakers separated, ready to quote or code. Files up to 10 minutes are free without an account; full interviews with speaker labels come with a free account, then you pay per minute – no subscription.
Speakers separated: each turn is labelled, so a one-hour interview reads as a conversation, not a wall of text.
Quote-ready: edit the text online, download as TXT, or as SRT/VTT with timestamps to find a passage in the recording.
Pay per hour, not per month: one interview or fifty – buy the minutes you need, they never expire.
99+ languages: the language is detected automatically.
Confidential: uploads are deleted after 30 days; delete them earlier yourself. No file is ever shared.
Free up to 10 minutes, no registration required.
Supported formats: MP3, M4A, WAV, MP4, MOV, OGG, FLAC · Prices
The transcript marks each change of speaker and labels the turns (Speaker 1, Speaker 2 …). You rename them in the editor. Speaker labels are included for registered users; guests get plain text.
For a clear recording it produces a near-verbatim transcript that needs a quick read-through rather than a re-listen. Check quotes against the recording before publishing – the timestamps in the SRT/VTT export take you straight to the passage.
Yes – 30% off every pack for non-profits, educators and researchers. Email contact@transcribe.mov from your organisation's address and you'll have your discount code within a day.
Up to 10 minutes without an account. With a free account you get 30 free minutes and can upload files up to 4 GB; after that you pay per minute – 120 minutes for $5.90, 300 for $10.90, 600 for $19.90.
Nobody but you. Files are processed automatically, never shared, deleted from our servers after 30 days, and you can delete a transcript and its file at any time from your account.
Transcribing a recording without multiple speakers? Transcribe a recording
Sign up now to start transcribing your audio and video files with high accuracy, secure and private processing, and industry-leading features like speaker diarization and subtitle generation.