Transcription & speech
Find the tool you need, then open it to get started.
All transcription & speech
Audio to text
Transcribe a short audio recording into text and timed JSON.
Audio translator
Transcribe a recording and translate its speech into text.
Meeting notes generator
Find decisions, action items and open questions in meeting notes.
PDF to audio
Read selectable PDF text aloud using a local system voice.
Podcast chapter generator
Draft podcast chapters using timestamps already in your transcript.
Pronunciation player
Hear words and phrases with an available browser voice.
Read aloud
Listen to written text with a browser voice.
Speaker diarization
Group non-overlapping voices using local ECAPA speaker embeddings.
Text to speech
Read text aloud with a browser voice.
Transcript editor
Import transcript JSON and local audio.
Transcript search
Search locally imported timestamped transcripts by phrase and semantic similarity.
Transcript summarizer
Summarize a transcript with supporting source quotations.
Video dubbing
Translate speech and replace the audio using an installed local system voice.
Video to text
Transcribe a video’s speech into text and timed JSON.
Voice cloning
Create a short English speech sample from a voice you are permitted to use.
Voice to text
Turn an uploaded voice recording into a transcript.
Voiceover timing
Estimate the speaking pace needed for a target duration.
No tools match that search. Try a shorter phrase or a file format.
Transcription, translation and speech generation
Transcription converts spoken audio into text. Translation changes the language of that text. Text-to-speech creates audio from written words. They can form one workflow, but each stage needs its own review: a transcription error can carry through into a translated subtitle or voiceover.
How to choose your next step
Use a clear recording where possible. Check names, numbers and overlapping speakers first. Keep the transcript editable so that corrected wording can flow into captions, summaries or narration.
Common questions about transcription & speech
Are speaker labels part of ordinary transcription?
Not necessarily. Transcription identifies words; diarization estimates which speaker spoke when. A label such as Speaker 1 does not establish a person’s identity. Overlapping voices can make both tasks harder.
Is reading text aloud the same as downloading generated speech?
No. Browser read-aloud uses voices available to the browser or operating system, and some voices use remote services. Creating a downloadable audio file needs a separate synthesis and export pipeline.