Export decisions · Buyer guide

AI Transcription Tools with SRT, VTT and Speaker Labels Compared

Direct answer

If the deliverable is a subtitle file, compare exports before transcription claims. Descript documents both SRT and VTT with an option to include speaker labels. TurboScribe advertises free SRT/VTT downloads for short uploaded files, but we could not confirm speaker-label preservation in each exported format. Otter Basic exports TXT only; SRT requires a paid plan, and VTT is not listed in its export documentation. For an interview handoff, keep an editable transcript as well as any timed captions. For a web video, confirm the player’s required subtitle format before uploading the recording. Word highlighting in an editor is not proof of word-level timestamps in a downloaded file. The comparison below separates file availability, free-plan limits and unresolved export details. Sources: Descript, TurboScribe, Otter.

AI Tool Finder Editorial Team · Sources checked October 7, 2026

Evidence: official documentation reviewed. Product promises remain vendor claims; our task recommendations are editorial interpretation. No account generation or export test was performed. The acceptance steps below are proposed checks, not results.

Compare the files, not just the editor

On a small screen, swipe the table horizontally. Checked October 7, 2026; documentation, not a completed export test.

What can you hand to the next person?
Tool / use caseSRT and VTTSpeakers and timingFree download / text files
Descript — edit media and captionsSRT: yes. VTT: yes. Subtitle controls“Show speakers” is documented for subtitle export; line/card controls affect caption segmentation. No word-level file guarantee.Free has 60 media minutes/month. Exact Free SRT/VTT entitlement is not stated in the reviewed export help: verify before committing. TXT/DOCX/CSV options not independently confirmed here. Plan table
TurboScribe — existing short filesSRT: yes. VTT: yes; free download advertised. Free captionsSpeaker recognition labels transcript sections. Label survival in each subtitle format and timestamp granularity: UNKNOWN.Free: 3 files/day, up to 30 minutes each, one at a time. DOCX/PDF/TXT also listed; CSV not listed. Export list
Otter — meeting transcript handoffSRT: paid. VTT: not documented in reviewed export list. Export helpSpeaker-name and timestamp export controls exist; confirm which apply to your chosen file. SRT carries caption timings.Basic: TXT only. Paid: DOCX/PDF/SRT. Removing branding is documented for TXT/clipboard, not as a universal switch. Plan and options

Three short task routes

Podcast → transcript → SRT

Finish the audio edit first. Generate captions against that final timeline, correct names, then export SRT. An earlier transcript can drift after cuts.

Interview → speakers → editable text

Correct speaker labels before export. Deliver an editable document plus the permitted recording; keep uncertain identities marked rather than guessing.

Existing video → captions → VTT

Choose a workflow that explicitly offers VTT. Confirm line breaks, cue timing and names in the target web player; do not stop at a successful download.

These are proposed workflows, not three tests performed by AI Tool Finder.

Inputs, plans and language boundaries

Descript: keep the edit and caption timeline together

Descript’s current pricing table lists 60 media minutes/month and a 1GB upload ceiling on Free. Hobbyist lists 10 media hours/month and 10GB files; its monthly price is US$24 per person, versus the lower annual-billing equivalent. The table lists 25 transcription languages, including English, Spanish, French and German. These limits are not the same as translation credits. Pricing and feature table.

Works when: you want to correct captions while editing the source media. Does not fit when: a free subtitle-download entitlement must be guaranteed before signup, or your source language is absent from the listed set. The separate subtitle export supports SRT/VTT and speaker labels; styled captions rendered into video are another deliverable. Subtitle export help.

Audio/video input is supported in the editor, but an exhaustive current codec list and an independent per-file duration cap were not verified here. Check the actual source container before uploading a large archive.

TurboScribe: a file-first option with a clearly bounded free route

The free captions tool accepts common audio/video formats including MP3, WAV, M4A, MP4 and MOV. It advertises SRT/VTT downloads and limits Free to three files daily, each at most 30 minutes, processed one at a time. Free caption tool.

The vendor lists 98+ transcription languages, separately from translation. Unlimited is advertised at US$120/year (US$10/month equivalent), with up to 10 hours or 5GB per file. Plans and language scope.

Works when: the recording already exists and you want a subtitle file rather than a meeting bot. Does not fit when: every speaker name must be proven to survive the first export without a check. The reviewed material confirms speaker recognition, but not that exact combination for every file type. Do not substitute the vendor’s accuracy headline for listening to your sample.

Otter: useful documents, but Basic is not a free SRT route

Basic lists 300 monthly minutes, 30 minutes per conversation and three lifetime audio/video imports. Pro lists ten monthly imports and 90-minute conversations; recurring meeting recording and imported-file allowances should not be conflated. Current plan limits.

Otter accepts audio such as MP3/WAV and video such as MP4/MOV. Import overview. English, Spanish, French, German, Japanese and Simplified Chinese are listed as transcription languages; asking Chat to translate is a separate operation. Language support.

Works when: a corrected meeting transcript is primary and paid SRT is enough. Does not fit when: you need free subtitle downloads, native VTT, or recurring imports beyond your plan. Individual export options and shared-file permissions still matter. Export controls.

What “timestamps and speaker labels” must mean in your brief

Write the timing requirement in the delivery brief. A transcript paragraph with a start time, a subtitle cue with a start and end, and a word-level alignment are different structures. A product can display all three internally without exporting all three.

For a subtitle file, ask whether a speaker name becomes visible caption text, a structured field or no output at all. If the recipient needs analysis by speaker, a timed text file may be less convenient than a document or structured export. None of the three reviewed export lists establishes a universal CSV speaker-and-word-timestamp delivery; do not promise one.

Original-language transcription and translated subtitles also need separate approval. A translated line can need different reading time, line breaks and terminology. A language count does not establish equivalent quality across languages, accents or overlapping voices.

A live meeting bot is not required to process an existing file. Upload only recordings you are permitted to share, and confirm retention and deletion rules for the specific workspace before sending sensitive interviews.

Inspect the file before accepting the delivery

  1. Choose a short, permissioned sample with two speakers, one proper name, a pause and one interruption. Record the language and source duration.
  2. Request the exact deliverables separately: editable transcript, SRT and/or VTT. State whether names must appear in captions.
  3. Correct the text and speaker names. Export after the final audio/video edit.
  4. Open the downloaded text file. Check cue start/end times, non-empty text, names, encoding and obvious malformed entries.
  5. Load it into the intended player or editor. Check the beginning, a middle speaker change and the ending for drift and readable line lengths.
  6. Keep the source, corrected transcript, final subtitle file and export settings together. Record any manual conversion as another step, not native support.

Proposed acceptance test only: no audio was uploaded, no subtitle file was generated and no output was timed by us for this article. The checklist is not an accuracy benchmark.

Why these three, and where the other guides fit

The shortlist covers an editor with explicit speaker-labelled subtitle controls, a short-file free download route and a meeting assistant with a clear paid subtitle gate. SpotScribe remains a separate podcast-link workflow; its existing profile does not establish the complete SRT/VTT-plus-speaker delivery contract. Adobe Podcast and ElevenLabs were not included without verifying that complete web-app workflow and entitlement. This is an evidence boundary, not a claim that they cannot export subtitles.

Use free transcription tools for allowance shopping, speaker diarization tools for recognition/deployment choices, and podcast post-production for editing the whole episode.

Next step: send the recipient a proposed delivery list before choosing the tool: “original-language transcript + corrected speaker names + VTT for this player.” Resolve any UNKNOWN in the table before committing a long recording.