{
  "title": "Audio to text",
  "human_url": "https://happytails.ai/en/speech-tools/audio-to-text/",
  "agent_url": "https://happytails.ai/agents/en/speech-tools/audio-to-text/",
  "language": "en",
  "markdown_url": "https://happytails.ai/en/speech-tools/audio-to-text.md",
  "content": "Audio to text\nTranscribe a short audio recording into text and timed JSON.\nUseful next steps\nTranscript editor\nImport transcript JSON and local audio.\nAuto subtitles\nGenerate draft SRT captions from recorded speech.\nTranscript summarizer\nSummarize a transcript with supporting source quotations.\nRemove background noise\nReduce steady background noise in an audio recording.\nHow to use Audio to text\nChoose a file in WAV, MP3, M4A, FLAC, MP4, MOV or WEBM.\nReview the settings, then choose Convert file.\nReview the result, then use Copy result or a download link when available.\nWhat to expect\nTranscribe a short audio recording into text and timed JSON. Runs with local model weights on this website’s server. No external AI API is used. Review the result for errors. Local Whisper tiny. Clips up to five minutes; review words and timing.",
  "content_format": "text/plain",
  "visibility": "public",
  "links": [
    {
      "name": "Transcript editor Import transcript JSON and local audio.",
      "url": "https://happytails.ai/en/speech-tools/transcript-editor/"
    },
    {
      "name": "Auto subtitles Generate draft SRT captions from recorded speech.",
      "url": "https://happytails.ai/en/subtitle-tools/auto-subtitles/"
    },
    {
      "name": "Transcript summarizer Summarize a transcript with supporting source quotations.",
      "url": "https://happytails.ai/en/speech-tools/transcript-summarizer/"
    },
    {
      "name": "Remove background noise Reduce steady background noise in an audio recording.",
      "url": "https://happytails.ai/en/media-tools/remove-background-noise/"
    }
  ],
  "tool": {
    "id": "audio-to-text",
    "category": "speech-tools",
    "built": true,
    "requirements": "Upload audio; select or detect language; edit timestamped transcript; export TXT and JSON; show uncertain passages and suppress silence hallucinations. Test accents, noise and long-file chunk boundaries.",
    "implementation": {
      "fields": [],
      "files": true,
      "accept": ".wav,.mp3,.m4a,.flac,.mp4,.mov,.webm",
      "multiple": false,
      "input": false,
      "processing": "Server required",
      "native": true,
      "ai": true,
      "maxFileBytes": 20971520,
      "note": "Runs with local model weights on this website’s server. No external AI API is used. Review the result for errors. Local Whisper tiny. Clips up to five minutes; review words and timing.",
      "experience": {
        "automatic": false,
        "single": false,
        "label": "Your text",
        "placeholder": "Type or paste your input…",
        "help": "Up to 1,000,000 characters. Review the settings before running.",
        "action": "Convert file"
      },
      "mode": "explicit",
      "interactive": false
    },
    "execution": "POST /api/v1/tools/audio-to-text/run",
    "api": {
      "id": "audio-to-text",
      "name": "Audio to text",
      "category": "speech-tools",
      "description": "Transcribe a short audio recording into text and timed JSON.",
      "notes": "Runs with local model weights on this website’s server. No external AI API is used. Review the result for errors. Local Whisper tiny. Clips up to five minutes; review words and timing.",
      "human_path": "/en/speech-tools/audio-to-text/",
      "method": "POST",
      "endpoint": "/api/v1/tools/audio-to-text/run",
      "schema_url": "/api/v1/tools/audio-to-text",
      "input_schema": {
        "type": "object",
        "additionalProperties": false,
        "properties": {
          "input": {
            "type": "string",
            "maxLength": 1000000,
            "default": "",
            "description": "Plain text input. File tools use files instead unless otherwise documented."
          },
          "options": {
            "type": "object",
            "additionalProperties": false,
            "properties": {}
          },
          "files": {
            "type": "array",
            "maxItems": 1,
            "items": {
              "type": "object",
              "required": [
                "name",
                "base64"
              ],
              "additionalProperties": false,
              "properties": {
                "name": {
                  "type": "string",
                  "maxLength": 200,
                  "description": "Filename only, no path."
                },
                "mime": {
                  "type": "string",
                  "maxLength": 150
                },
                "bytes": {
                  "type": "integer",
                  "minimum": 0,
                  "description": "Optional decoded byte count; must match content if supplied."
                },
                "base64": {
                  "type": "string",
                  "contentEncoding": "base64",
                  "description": "File bytes as padded base64. Use the tool-specific limits.files_bytes value for the total decoded input size."
                }
              }
            }
          }
        }
      },
      "output_schema": {
        "type": "object",
        "required": [
          "tool",
          "result"
        ],
        "properties": {
          "tool": {
            "type": "string"
          },
          "result": {
            "type": "object",
            "required": [
              "text",
              "files"
            ],
            "properties": {
              "text": {
                "type": [
                  "string",
                  "null"
                ]
              },
              "data": {
                "description": "Parsed JSON when the textual result is JSON."
              },
              "name": {
                "type": "string"
              },
              "mime": {
                "type": "string"
              },
              "files": {
                "type": "array",
                "items": {
                  "type": "object",
                  "required": [
                    "name",
                    "mime",
                    "bytes",
                    "base64"
                  ],
                  "properties": {
                    "name": {
                      "type": "string"
                    },
                    "mime": {
                      "type": "string"
                    },
                    "bytes": {
                      "type": "integer"
                    },
                    "base64": {
                      "type": "string",
                      "contentEncoding": "base64"
                    }
                  }
                }
              }
            }
          }
        }
      },
      "processing": "Server; self-hosted native engine, no external conversion service",
      "limits": {
        "request_bytes": 31457280,
        "input_characters": 1000000,
        "files_bytes": 20971520,
        "max_files": 1,
        "timeout_seconds": 90
      }
    },
    "agent_processing": "Server; self-hosted native engine, no external conversion service",
    "browser_agent": {
      "supported": true,
      "name": "run_current_tool",
      "discovery": "WebMCP on the human page in a supporting browser",
      "verification": "See audit/API-AUDIT.md; availability is not verification"
    }
  },
  "markdown": "# Audio to text\n\n- Human page: https://happytails.ai/en/speech-tools/audio-to-text/\n- Markdown: https://happytails.ai/en/speech-tools/audio-to-text.md\n- Structured JSON: https://happytails.ai/agents/en/speech-tools/audio-to-text/index.json\n- Visibility: public\n\n## Tool capabilities\n\n| Property | Value |\n| --- | --- |\n| Tool ID | `audio-to-text` |\n| Category | speech-tools |\n| Implementation | Registered implementation. Registration alone is not a production-quality guarantee. |\n| Browser execution | Use the visible action or interactive controls. |\n| Browser processing | Server required |\n| HTTP API | /api/v1/tools/audio-to-text/run |\n| Browser-agent interface | run_current_tool |\n\n### Implementation notes\n\nRuns with local model weights on this website’s server. No external AI API is used. Review the result for errors. Local Whisper tiny. Clips up to five minutes; review words and timing.\n\n### Inputs and settings\n\nNo free-text input is required. Use the file inputs and/or settings below.\n\nAccepted browser files: `.wav,.mp3,.m4a,.flac,.mp4,.mov,.webm`. Choose one file.\n\nBrowser file-input limit: 20971520 bytes in total. HTTP limits can differ.\n\nNo additional settings are declared for this tool.\n\n### HTTP contract\n\nAPI calls process supplied data on the server. The website origin is `https://happytails.ai`. Server API hosting is not connected on the public website yet. Use your local server origin to test these requests.\n\n```http\nPOST https://happytails.ai/api/v1/tools/audio-to-text/run\nContent-Type: application/json\n```\n\n[Detailed operation reference](https://happytails.ai/en/api/tools/audio-to-text.md) · [Shared API guide](https://happytails.ai/en/api.md)\n\n#### Limits\n\n| Limit | Value |\n| --- | --- |\n| request_bytes | 31457280 |\n| input_characters | 1000000 |\n| files_bytes | 20971520 |\n| max_files | 1 |\n| timeout_seconds | 90 |\n\n#### Complete input schema\n\n```json\n{\n  \"type\": \"object\",\n  \"additionalProperties\": false,\n  \"properties\": {\n    \"input\": {\n      \"type\": \"string\",\n      \"maxLength\": 1000000,\n      \"default\": \"\",\n      \"description\": \"Plain text input. File tools use files instead unless otherwise documented.\"\n    },\n    \"options\": {\n      \"type\": \"object\",\n      \"additionalProperties\": false,\n      \"properties\": {}\n    },\n    \"files\": {\n      \"type\": \"array\",\n      \"maxItems\": 1,\n      \"items\": {\n        \"type\": \"object\",\n        \"required\": [\n          \"name\",\n          \"base64\"\n        ],\n        \"additionalProperties\": false,\n        \"properties\": {\n          \"name\": {\n            \"type\": \"string\",\n            \"maxLength\": 200,\n            \"description\": \"Filename only, no path.\"\n          },\n          \"mime\": {\n            \"type\": \"string\",\n            \"maxLength\": 150\n          },\n          \"bytes\": {\n            \"type\": \"integer\",\n            \"minimum\": 0,\n            \"description\": \"Optional decoded byte count; must match content if supplied.\"\n          },\n          \"base64\": {\n            \"type\": \"string\",\n            \"contentEncoding\": \"base64\",\n            \"description\": \"File bytes as padded base64. Use the tool-specific limits.files_bytes value for the total decoded input size.\"\n          }\n        }\n      }\n    }\n  }\n}\n```\n\n#### Complete output schema\n\n```json\n{\n  \"type\": \"object\",\n  \"required\": [\n    \"tool\",\n    \"result\"\n  ],\n  \"properties\": {\n    \"tool\": {\n      \"type\": \"string\"\n    },\n    \"result\": {\n      \"type\": \"object\",\n      \"required\": [\n        \"text\",\n        \"files\"\n      ],\n      \"properties\": {\n        \"text\": {\n          \"type\": [\n            \"string\",\n            \"null\"\n          ]\n        },\n        \"data\": {\n          \"description\": \"Parsed JSON when the textual result is JSON.\"\n        },\n        \"name\": {\n          \"type\": \"string\"\n        },\n        \"mime\": {\n          \"type\": \"string\"\n        },\n        \"files\": {\n          \"type\": \"array\",\n          \"items\": {\n            \"type\": \"object\",\n            \"required\": [\n              \"name\",\n              \"mime\",\n              \"bytes\",\n              \"base64\"\n            ],\n            \"properties\": {\n              \"name\": {\n                \"type\": \"string\"\n              },\n              \"mime\": {\n                \"type\": \"string\"\n              },\n              \"bytes\": {\n                \"type\": \"integer\"\n              },\n              \"base64\": {\n                \"type\": \"string\",\n                \"contentEncoding\": \"base64\"\n              }\n            }\n          }\n        }\n      }\n    }\n  }\n}\n```\n\n#### Error and retry handling\n\nFailures return a non-200 status and `error.code` plus `error.message`. Correct invalid input before retrying; retry server-busy responses with bounded backoff. Returned files contain Base64 data, not persistent download URLs. Decode and inspect the output before treating conversion as successful. See the shared API guide for the complete status and timeout rules.\n\n### Browser-agent workflow\n\n1. Open the human page in a browser that supports WebMCP.\n2. Discover the interface exposed by that page; use its actual schema.\n3. Select files on the page first when the tool requires files.\n4. Call `run_current_tool` with valid input and settings.\n5. Inspect the returned result or error and the displayed output. Retrieve files from the displayed download links.\n\nCalls do not automatically copy or download results. Browser permissions still apply. Interface availability alone is not proof of successful execution.\n\n### Planning specification\n\nUpload audio; select or detect language; edit timestamped transcript; export TXT and JSON; show uncertain passages and suppress silence hallucinations. Test accents, noise and long-file chunk boundaries.\n\nThis describes intended requirements. Use the implementation notes, interface schema and observed output to determine current support.\n\n### Complete machine-readable capability record\n\n```json\n{\n  \"id\": \"audio-to-text\",\n  \"category\": \"speech-tools\",\n  \"built\": true,\n  \"requirements\": \"Upload audio; select or detect language; edit timestamped transcript; export TXT and JSON; show uncertain passages and suppress silence hallucinations. Test accents, noise and long-file chunk boundaries.\",\n  \"implementation\": {\n    \"fields\": [],\n    \"files\": true,\n    \"accept\": \".wav,.mp3,.m4a,.flac,.mp4,.mov,.webm\",\n    \"multiple\": false,\n    \"input\": false,\n    \"processing\": \"Server required\",\n    \"native\": true,\n    \"ai\": true,\n    \"maxFileBytes\": 20971520,\n    \"note\": \"Runs with local model weights on this website’s server. No external AI API is used. Review the result for errors. Local Whisper tiny. Clips up to five minutes; review words and timing.\",\n    \"experience\": {\n      \"automatic\": false,\n      \"single\": false,\n      \"label\": \"Your text\",\n      \"placeholder\": \"Type or paste your input…\",\n      \"help\": \"Up to 1,000,000 characters. Review the settings before running.\",\n      \"action\": \"Convert file\"\n    },\n    \"mode\": \"explicit\",\n    \"interactive\": false\n  },\n  \"execution\": \"POST /api/v1/tools/audio-to-text/run\",\n  \"api\": {\n    \"id\": \"audio-to-text\",\n    \"name\": \"Audio to text\",\n    \"category\": \"speech-tools\",\n    \"description\": \"Transcribe a short audio recording into text and timed JSON.\",\n    \"notes\": \"Runs with local model weights on this website’s server. No external AI API is used. Review the result for errors. Local Whisper tiny. Clips up to five minutes; review words and timing.\",\n    \"human_path\": \"/en/speech-tools/audio-to-text/\",\n    \"method\": \"POST\",\n    \"endpoint\": \"/api/v1/tools/audio-to-text/run\",\n    \"schema_url\": \"/api/v1/tools/audio-to-text\",\n    \"input_schema\": {\n      \"type\": \"object\",\n      \"additionalProperties\": false,\n      \"properties\": {\n        \"input\": {\n          \"type\": \"string\",\n          \"maxLength\": 1000000,\n          \"default\": \"\",\n          \"description\": \"Plain text input. File tools use files instead unless otherwise documented.\"\n        },\n        \"options\": {\n          \"type\": \"object\",\n          \"additionalProperties\": false,\n          \"properties\": {}\n        },\n        \"files\": {\n          \"type\": \"array\",\n          \"maxItems\": 1,\n          \"items\": {\n            \"type\": \"object\",\n            \"required\": [\n              \"name\",\n              \"base64\"\n            ],\n            \"additionalProperties\": false,\n            \"properties\": {\n              \"name\": {\n                \"type\": \"string\",\n                \"maxLength\": 200,\n                \"description\": \"Filename only, no path.\"\n              },\n              \"mime\": {\n                \"type\": \"string\",\n                \"maxLength\": 150\n              },\n              \"bytes\": {\n                \"type\": \"integer\",\n                \"minimum\": 0,\n                \"description\": \"Optional decoded byte count; must match content if supplied.\"\n              },\n              \"base64\": {\n                \"type\": \"string\",\n                \"contentEncoding\": \"base64\",\n                \"description\": \"File bytes as padded base64. Use the tool-specific limits.files_bytes value for the total decoded input size.\"\n              }\n            }\n          }\n        }\n      }\n    },\n    \"output_schema\": {\n      \"type\": \"object\",\n      \"required\": [\n        \"tool\",\n        \"result\"\n      ],\n      \"properties\": {\n        \"tool\": {\n          \"type\": \"string\"\n        },\n        \"result\": {\n          \"type\": \"object\",\n          \"required\": [\n            \"text\",\n            \"files\"\n          ],\n          \"properties\": {\n            \"text\": {\n              \"type\": [\n                \"string\",\n                \"null\"\n              ]\n            },\n            \"data\": {\n              \"description\": \"Parsed JSON when the textual result is JSON.\"\n            },\n            \"name\": {\n              \"type\": \"string\"\n            },\n            \"mime\": {\n              \"type\": \"string\"\n            },\n            \"files\": {\n              \"type\": \"array\",\n              \"items\": {\n                \"type\": \"object\",\n                \"required\": [\n                  \"name\",\n                  \"mime\",\n                  \"bytes\",\n                  \"base64\"\n                ],\n                \"properties\": {\n                  \"name\": {\n                    \"type\": \"string\"\n                  },\n                  \"mime\": {\n                    \"type\": \"string\"\n                  },\n                  \"bytes\": {\n                    \"type\": \"integer\"\n                  },\n                  \"base64\": {\n                    \"type\": \"string\",\n                    \"contentEncoding\": \"base64\"\n                  }\n                }\n              }\n            }\n          }\n        }\n      }\n    },\n    \"processing\": \"Server; self-hosted native engine, no external conversion service\",\n    \"limits\": {\n      \"request_bytes\": 31457280,\n      \"input_characters\": 1000000,\n      \"files_bytes\": 20971520,\n      \"max_files\": 1,\n      \"timeout_seconds\": 90\n    }\n  },\n  \"agent_processing\": \"Server; self-hosted native engine, no external conversion service\",\n  \"browser_agent\": {\n    \"supported\": true,\n    \"name\": \"run_current_tool\",\n    \"discovery\": \"WebMCP on the human page in a supporting browser\",\n    \"verification\": \"See audit/API-AUDIT.md; availability is not verification\"\n  }\n}\n```\n\n## Page guide\n\nTranscribe a short audio recording into text and timed JSON.\n\n## Useful next steps\n\n- [Transcript editor](https://happytails.ai/en/speech-tools/transcript-editor/): Import transcript JSON and local audio. ([Markdown](https://happytails.ai/en/speech-tools/transcript-editor.md))\n\n- [Auto subtitles](https://happytails.ai/en/subtitle-tools/auto-subtitles/): Generate draft SRT captions from recorded speech. ([Markdown](https://happytails.ai/en/subtitle-tools/auto-subtitles.md))\n\n- [Transcript summarizer](https://happytails.ai/en/speech-tools/transcript-summarizer/): Summarize a transcript with supporting source quotations. ([Markdown](https://happytails.ai/en/speech-tools/transcript-summarizer.md))\n\n- [Remove background noise](https://happytails.ai/en/media-tools/remove-background-noise/): Reduce steady background noise in an audio recording. ([Markdown](https://happytails.ai/en/media-tools/remove-background-noise.md))\n\n## How to use Audio to text\n\n1. Choose a file in WAV, MP3, M4A, FLAC, MP4, MOV or WEBM.\n2. Review the settings, then choose Convert file.\n3. Review the result, then use Copy result or a download link when available.\n\n## What to expect\n\nTranscribe a short audio recording into text and timed JSON. Runs with local model weights on this website’s server. No external AI API is used. Review the result for errors. Local Whisper tiny. Clips up to five minutes; review words and timing.\n"
}