# Transcription & speech

- Human page: https://happytails.ai/en/speech-tools/
- Markdown: https://happytails.ai/en/speech-tools.md
- Structured JSON: https://happytails.ai/agents/en/speech-tools/index.json
- Visibility: public

## Page guide

Find the tool you need, then open it to get started.

## All transcription & speech

- [Audio to text](https://happytails.ai/en/speech-tools/audio-to-text/): Transcribe a short audio recording into text and timed JSON. ([Markdown](https://happytails.ai/en/speech-tools/audio-to-text.md))

- [Audio translator](https://happytails.ai/en/speech-tools/audio-translator/): Transcribe a recording and translate its speech into text. ([Markdown](https://happytails.ai/en/speech-tools/audio-translator.md))

- [Meeting notes generator](https://happytails.ai/en/speech-tools/meeting-notes-generator/): Find decisions, action items and open questions in meeting notes. ([Markdown](https://happytails.ai/en/speech-tools/meeting-notes-generator.md))

- [PDF to audio](https://happytails.ai/en/speech-tools/pdf-to-audio/): Read selectable PDF text aloud using a local system voice. ([Markdown](https://happytails.ai/en/speech-tools/pdf-to-audio.md))

- [Podcast chapter generator](https://happytails.ai/en/speech-tools/podcast-chapter-generator/): Draft podcast chapters using timestamps already in your transcript. ([Markdown](https://happytails.ai/en/speech-tools/podcast-chapter-generator.md))

- [Pronunciation player](https://happytails.ai/en/speech-tools/pronunciation-player/): Hear words and phrases with an available browser voice. ([Markdown](https://happytails.ai/en/speech-tools/pronunciation-player.md))

- [Read aloud](https://happytails.ai/en/speech-tools/read-aloud/): Listen to written text with a browser voice. ([Markdown](https://happytails.ai/en/speech-tools/read-aloud.md))

- [Speaker diarization](https://happytails.ai/en/speech-tools/speaker-diarization/): Group non-overlapping voices using local ECAPA speaker embeddings. ([Markdown](https://happytails.ai/en/speech-tools/speaker-diarization.md))

- [Text to speech](https://happytails.ai/en/speech-tools/text-to-speech/): Read text aloud with a browser voice. ([Markdown](https://happytails.ai/en/speech-tools/text-to-speech.md))

- [Transcript editor](https://happytails.ai/en/speech-tools/transcript-editor/): Import transcript JSON and local audio. ([Markdown](https://happytails.ai/en/speech-tools/transcript-editor.md))

- [Transcript search](https://happytails.ai/en/speech-tools/transcript-search/): Search locally imported timestamped transcripts by phrase and semantic similarity. ([Markdown](https://happytails.ai/en/speech-tools/transcript-search.md))

- [Transcript summarizer](https://happytails.ai/en/speech-tools/transcript-summarizer/): Summarize a transcript with supporting source quotations. ([Markdown](https://happytails.ai/en/speech-tools/transcript-summarizer.md))

- [Video dubbing](https://happytails.ai/en/speech-tools/video-dubbing/): Translate speech and replace the audio using an installed local system voice. ([Markdown](https://happytails.ai/en/speech-tools/video-dubbing.md))

- [Video to text](https://happytails.ai/en/speech-tools/video-to-text/): Transcribe a video’s speech into text and timed JSON. ([Markdown](https://happytails.ai/en/speech-tools/video-to-text.md))

- [Voice cloning](https://happytails.ai/en/speech-tools/voice-cloning/): Create a short English speech sample from a voice you are permitted to use. ([Markdown](https://happytails.ai/en/speech-tools/voice-cloning.md))

- [Voice to text](https://happytails.ai/en/speech-tools/voice-to-text/): Turn an uploaded voice recording into a transcript. ([Markdown](https://happytails.ai/en/speech-tools/voice-to-text.md))

- [Voiceover timing](https://happytails.ai/en/speech-tools/voiceover-timing/): Estimate the speaking pace needed for a target duration. ([Markdown](https://happytails.ai/en/speech-tools/voiceover-timing.md))

## Transcription, translation and speech generation

Transcription converts spoken audio into text. Translation changes the language of that text. Text-to-speech creates audio from written words. They can form one workflow, but each stage needs its own review: a transcription error can carry through into a translated subtitle or voiceover.

## How to choose your next step

Use a clear recording where possible. Check names, numbers and overlapping speakers first. Keep the transcript editable so that corrected wording can flow into captions, summaries or narration.

## Common questions about transcription & speech

### Are speaker labels part of ordinary transcription?

Not necessarily. Transcription identifies words; diarization estimates which speaker spoke when. A label such as Speaker 1 does not establish a person’s identity. Overlapping voices can make both tasks harder.

### Is reading text aloud the same as downloading generated speech?

No. Browser read-aloud uses voices available to the browser or operating system, and some voices use remote services. Creating a downloadable audio file needs a separate synthesis and export pipeline.
